Advanced Evaluation
guanyang/open-agent-hub
This skill should be used for advanced LLM evaluation: LLM-as-judge systems, direct scoring, pairwise comparison, rubric calibration, evaluator bias mitigation, confidence scoring, and automated…
Decomposed, multi-criteria metric design for LLM pipelines. An agent skill from agentsope/SkillAlchemy.
$ npx skills add agentsope/SkillAlchemy --skill agentsop-metric-design -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install agentsope/SkillAlchemy agentsop-metric-design --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/agentsope/SkillAlchemy.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/agentsop-metric-design .claude/skills/agentsop-metric-design && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "agentsop-metric-design" agent skill from https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-metric-design into .claude/skills/agentsop-metric-design/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agentsop-metric-design", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-metric-designType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add agentsope/SkillAlchemy --skill agentsop-metric-design -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install agentsope/SkillAlchemy agentsop-metric-design --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/agentsope/SkillAlchemy.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/agentsop-metric-design .agents/skills/agentsop-metric-design && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "agentsop-metric-design" agent skill from https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-metric-design into .agents/skills/agentsop-metric-design/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agentsop-metric-design", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add agentsope/SkillAlchemy --skill agentsop-metric-design -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install agentsope/SkillAlchemy agentsop-metric-design --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/agentsope/SkillAlchemy.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/agentsop-metric-design .cursor/skills/agentsop-metric-design && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "agentsop-metric-design" agent skill from https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-metric-design into .cursor/skills/agentsop-metric-design/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agentsop-metric-design", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/agentsope/SkillAlchemy.git --path skills/agentsop-metric-design--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add agentsope/SkillAlchemy --skill agentsop-metric-design -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install agentsope/SkillAlchemy agentsop-metric-design --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/agentsope/SkillAlchemy.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/agentsop-metric-design .gemini/skills/agentsop-metric-design && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "agentsop-metric-design" agent skill from https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-metric-design into .gemini/skills/agentsop-metric-design/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agentsop-metric-design", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install agentsope/SkillAlchemy agentsop-metric-designInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add agentsope/SkillAlchemy --skill agentsop-metric-design -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/agentsope/SkillAlchemy.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/agentsop-metric-design .github/skills/agentsop-metric-design && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "agentsop-metric-design" agent skill from https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-metric-design into .github/skills/agentsop-metric-design/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agentsop-metric-design", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add agentsope/SkillAlchemy --skill agentsop-metric-design -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install agentsope/SkillAlchemy agentsop-metric-design --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/agentsope/SkillAlchemy.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/agentsop-metric-design .opencode/skills/agentsop-metric-design && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "agentsop-metric-design" agent skill from https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-metric-design into .opencode/skills/agentsop-metric-design/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agentsop-metric-design", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
agentsop-metric-designDecomposed, multi-criteria metric design for LLM pipelines. An agent skill from agentsope/SkillAlchemy.
Agentsop Metric Design is an agent skill from agentsope/SkillAlchemy. Decomposed, multi-criteria metric design for LLM pipelines. The metric IS the model — change the metric and the optimizer changes behavior. Decompose by default; bool during compile, float during eval; calibrate against human; mitigate judge bias. Search keywords: LLM-as-judge, llm as judge, eval metric, evaluation score, scoring function, rubric, RAGAS, G-Eval, judge bias, verbosity bias, how to evaluate LLM output.
Its SKILL.md is about 6.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including reference files (for example `README.md`, `intermediate/operation_candidates.json` and `references/R1-source-evidence.md`).
It sits in AI & LLM Engineering, covering LLM evaluation, Quizzes and assessments and Operations and SOPs. The repository describes itself as: From thought to skill. From signal to structure. The licence is MIT.
7 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit d0f0355. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are python).
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Agentsop Metric Design loads about 6.4k tokens when it runs, and up to ~11k if it reads all its reference files. Until then it costs about 111 tokens; SKILL.md has 2,651 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from agentsope/SkillAlchemy at commit d0f0355, republished under its MIT licence (© agentsope). 2,651 words, ~6,414 tokens.
.claude/skills/agentsop-metric-design/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub."It's unproductive to launch optimization runs using a poorly designed program or a bad metric." — DSPy core team [dspy.ai/learn/optimization/overview/]
"LLM judges exhibit self-preference, recency, rubric-order, score-ID, and length biases." — Synthesized from [arxiv.org/pdf/2506.02592, arxiv.org/pdf/2509.26072]
This is a tool skill. It produces a metric function (and a calibration receipt) that other skills consume — DSPy compilers (MIPROv2 / GEPA / BootstrapFewShot), LlamaIndex FaithfulnessEvaluator/RelevancyEvaluator/RetrieverEvaluator, LangGraph eval judges, RAGAS, TruLens. The metric is the optimization target. Get it wrong and every downstream optimizer is theatre.
Activate when any of these is true:
metric(example, pred) -> bool|float. The metric drives gradient-free search; bias propagates into the artifact.OP-10 EvalLoop).Do NOT activate for:
dspy-sop will refuse to compile anyway; this skill cannot help).A DSPy/GEPA/MIPRO optimizer is a black-box search that maximizes metric(pred, example). Whatever the metric rewards, the compiled prompt will produce. If the metric prefers verbose, hedged answers (which an LLM judge will, by default — judges over-prefer length [arxiv.org/pdf/2506.02592]), the optimizer will produce verbose, hedged answers. A bad metric beats a good optimizer every time.
Three corollaries:
trace argument. Compile-time bool prevents the optimizer from chasing noise in the middle of the distribution; eval-time float gives gradient for reporting and debugging. [dspy.ai/learn/evaluation/metrics/]| Bias | What it does | Source |
|---|---|---|
| Length bias | Judges prefer longer answers regardless of quality | [arxiv.org/pdf/2506.02592] |
| Self-preference | A judge from family X prefers outputs from family X | [arxiv.org/pdf/2506.02592] |
| Recency / position | Last option in a pairwise rated higher | [arxiv.org/pdf/2509.26072] |
| Rubric-order | Criteria listed first weighted more | [arxiv.org/pdf/2509.26072] |
| Score-ID | "5" and "10" anchor differently across rubrics | [arxiv.org/pdf/2509.26072] |
| Provenance | Knowing the source model biases the rating | [arxiv.org/pdf/2506.02592] |
Full catalog in references/R2-judge-bias-catalog.md.
0. Confirm activation criteria (§1)
1. DECOMPOSE: list orthogonal criteria
2. CHOICE: bool vs float per criterion
3. WRITE: implement sub-judges + aggregator
4. BIAS-TEST: probe for length, position, self-preference
5. CALIBRATE: ≥20 human spot-checks; agreement ≥80%
6. SHIP: hand metric to optimizer / eval loopWrite down the criteria the user actually cares about. For an open-ended Q&A:
For RAG specifically, copy LlamaIndex's triad: Faithfulness (answer entailed by context), Relevancy (answer addresses query), Retriever quality (MRR / hit-rate on labeled QA pairs). See references/R1-source-evidence.md.
Output of this stage: a checklist of 3–6 sub-criteria, each with type (bool / scalar) and aggregation rule (AND for hard gates, weighted sum for soft).
Per [dspy.ai/learn/evaluation/metrics/]:
| Criterion shape | Return | Why |
|---|---|---|
| Hard requirement (factual, schema-valid, no PII) | bool | Optimizer should reject anything that fails, not partially credit |
| Soft preference (concise, fluent, on-tone) | float ∈ [0,1] | Optimizer benefits from gradient |
| Length | explicit scalar penalty, not LLM-rated | LLM judges over-prefer length; bake the penalty in deterministically |
Universal rule (DSPy OP-030): wrap the whole metric so it returns bool when trace is not None (compile mode) and float otherwise (eval mode). Same function, two semantics.
import dspy
class Assess(dspy.Signature):
"""Assess a single binary criterion. Reply yes/no."""
text_to_assess = dspy.InputField()
assessment_question = dspy.InputField()
assessment_answer: bool = dspy.OutputField()
def metric(example, pred, trace=None):
# Hard gates (bool sub-judges)
factual = dspy.Predict(Assess)(text_to_assess=pred.answer,
assessment_question="Is every factual claim supported by the context?").assessment_answer
on_topic = dspy.Predict(Assess)(text_to_assess=pred.answer,
assessment_question=f"Does this address the question: '{example.question}'?").assessment_answer
non_hedging = dspy.Predict(Assess)(text_to_assess=pred.answer,
assessment_question="Does this answer commit to a position (no 'it depends' / 'as an AI' hedging)?").assessment_answer
# Length penalty (deterministic — DO NOT delegate to judge)
over_budget = len(pred.answer.split()) > 150
length_score = 1.0 if not over_budget else max(0.0, 1.0 - (len(pred.answer.split()) - 150) / 150)
if trace is not None: # compile mode → strict bool
return factual and on_topic and non_hedging and not over_budget
# eval mode → float for reporting
return (factual + on_topic + non_hedging) / 3.0 * length_scoreRules:
Before calibration, run these probes on a 20-example dev set:
Document failures in references/R2-judge-bias-catalog.md for this project.
This is non-negotiable. Per DSPy Case C: "Never compile against a metric you haven't human-validated on ≥ 20 spot-checks." Garbage metric → garbage compiled program.
Hand off:
{n_spot_checks, per_judge_agreement, bias_probe_results, judge_model_id, task_model_id, date}. Stored next to the compiled artifact. Required for any later audit.dspy.GEPA for sample efficiency; otherwise MIPROv2.dspy.Predict(Assess) call per criterion. Aggregate with AND (hard) + weighted sum (soft).teleprompt.compile) and dspy.Evaluate.if trace is not None: return bool_aggregate; else: return float_aggregate.length_score = clip(1 - max(0, len - budget) / budget, 0, 1) multiplied into the final float. Do NOT delegate length judgment to the LLM.FaithfulnessEvaluator (answer entailed by retrieved context), RelevancyEvaluator (answer addresses query), RetrieverEvaluator(["mrr","hit_rate"]) (does retrieval pull the gold passage). Gate every PR.dspy.Prediction(score=float, feedback=str) from the metric. Pair with dspy.GEPA optimizer (not MIPROv2).{metric_version, n_spot_checks, per_judge_agreement, bias_probe_results, judge_model_id, task_model_id, date, criteria_list}.save(save_program=True) analog.困境: Team compiled an open-ended customer-support response program with MIPROv2(metric=llm_judge, auto="medium"). Auto-eval score climbed from 0.62 to 0.84. Production rolled out. Users complained answers were too long. Auditor traced it: the LLM judge gave +0.18 to answers >120 words across the eval set, including some objectively wrong ones. The optimizer faithfully chased that signal.
约束: Re-compile costs ~$30. Cannot re-label dataset. Cannot change judge model (vendor approval cycle).
决策步骤:
OP-M03) and decompose the holistic judge into factual / on-topic / non-hedging sub-judges (OP-M01). Length is now deterministic, not judge-rated.OP-M05). Agreement: was 64%, now 87%.结果: The "score regression" (0.84 → 0.79) was misleading — the old metric was the bug. Real quality improved because the metric finally matched the user.
可提取的操作: OP-M01 DecomposeMultiCriteria, OP-M03 ExplicitLengthPenalty, OP-M05 HumanSpotCheck20. Lesson: when a compiled program is "good on metric, bad in production", the metric is wrong. Don't tune harder — fix the metric.
困境: RAG pipeline reports Faithfulness = 1.0 across 200 eval QA pairs. Users complain answers don't address their questions. Team is confused: "but we're 100% faithful…"
约束: Cannot change evaluator vendor. Existing eval suite has only FaithfulnessEvaluator.
决策步骤:
FaithfulnessEvaluator only asks "is the answer entailed by retrieved context?" An answer of "The context discusses Q3 financials" is faithful to the context but does not answer "What was Q3 revenue?".RelevancyEvaluator (does the answer address the question?) and RetrieverEvaluator(["mrr", "hit_rate"]) (is the gold passage even retrieved?). [LlamaIndex OP-10]. This is OP-M06.结果: A single metric (faithfulness) created a blind spot. Decomposition exposed the actual failure (synthesizer hedging) which no single number could surface.
可提取的操作: OP-M01, OP-M06 RAGFaithfulnessRelevancyContext. Lesson: a single RAG metric, however precise, is structurally incomplete. Triad is the minimum.
困境: A vendor-tied cekura-metric-design skill (11 installs) wraps a single proprietary judge. Convenient, but: (a) cannot inspect the rubric, (b) cannot swap judge family, (c) cannot add length penalty, (d) calibration receipt does not include bias probes. Optimizer is chasing the vendor judge's blind spots.
决策步骤:
OP-M01). Add length penalty (OP-M03), non-hedging sub-judge, cross-family fact-check (OP-M04).OP-M05), not the vendor's alone.结果: Vendor metrics buy convenience but lose audit and bias control. Open, decomposed metrics dominate for any pipeline that will be optimized against.
可提取的操作: OP-M01, OP-M03, OP-M04, OP-M10. Lesson: a vendor-tied holistic metric is a managed optimizer target you cannot inspect. Decompose around it.
| # | Anti-pattern | Why it's wrong | Fix |
|---|---|---|---|
| AP-1 | Single holistic LLM-as-judge ("rate 1-5") as optimizer metric | Conflates orthogonal axes; inherits all judge biases; optimizer chases noise | Decompose into yes/no sub-judges (OP-M01) |
| AP-2 | No length penalty in a generation metric | Judges over-prefer length; optimizer learns to be verbose | Deterministic length term (OP-M03) |
| AP-3 | Same family for judge and task model | Self-preference bias inflates score 5–15pp | Cross-family judge (OP-M04) |
| AP-4 | Skipping human calibration | Optimizer learns the metric's bias, not the task | ≥20 spot-checks (OP-M05) |
| AP-5 | Throwing optimizers at a stalled compile | If metric is bad, no optimizer can save it | Return to metric design |
| AP-6 | Single-axis RAG metric (faithfulness only) | Misses query-relevance and retrieval-quality | Triad (OP-M06) |
| AP-7 | LLM-rated length ("is this concise?" as a judge call) | Inherits length bias; deterministic is free | Length is len(tokens), not a judge call |
| AP-8 | Hand-edited metric mid-experiment without re-calibration | Past compile artifacts now have invalid calibration receipts | Bump metric version; re-calibrate |
| AP-9 | Float metric in compile mode | Optimizer chases sub-percent noise; overfit | bool in compile, float in eval (OP-M02) |
| AP-10 | Compiling without a metric at all | Without a metric DSPy degenerates to verbose prompting | Refuse the compile (DSPy OP-024) |
scientific-critical-thinking for problem definition first.simpo / trl-fine-tuning skills, not this one.| Concept | DSPy | LlamaIndex | RAGAS | TruLens | This skill |
|---|---|---|---|---|---|
| Metric type | def metric(ex, pred, trace=None) -> float|bool | EvaluationResult from BaseEvaluator | Metric class (Faithfulness, AnswerRelevancy, etc.) | Feedback function | metric function + calibration receipt |
| Multi-criteria | sub-judges via dspy.Predict(Assess) | stack of FaithfulnessEvaluator + RelevancyEvaluator + RetrieverEvaluator | metric ensemble | Feedback per criterion | OP-M01 DecomposeMultiCriteria |
| Mode switch (compile vs eval) | trace is not None returns bool | N/A (eval only) | N/A | N/A | OP-M02 |
| Faithfulness | sub-judge "every claim supported?" | FaithfulnessEvaluator (entailment-based) | Faithfulness metric (claim-by-claim NLI) | groundedness feedback | sub-judge or pre-built |
| Relevancy | sub-judge "addresses question?" | RelevancyEvaluator | AnswerRelevancy (embedding-based) | answer-relevance feedback | sub-judge |
| Retrieval quality | LlamaIndex RetrieverEvaluator MRR/hit-rate | RetrieverEvaluator(["mrr","hit_rate"]) | ContextPrecision, ContextRecall | context-relevance feedback | pre-built (use LlamaIndex / RAGAS) |
| Length penalty | manual scalar in metric | manual | manual | manual | OP-M03 (deterministic) |
| Textual feedback | dspy.Prediction(score, feedback) → GEPA | N/A | rationale strings | feedback rationale | OP-M08 |
| Cross-family judge | swap judge_lm in metric | swap service_context.llm of evaluator | swap evaluator LLM | swap feedback LLM | OP-M04 (always) |
Combination patterns:
FaithfulnessEvaluator + RelevancyEvaluator as the sub-judges inside a DSPy metric; aggregate to bool/float per OP-M02. (JetBlue/Databricks pattern.)Faithfulness / AnswerRelevancy / ContextPrecision as sub-scalars; pass into a DSPy metric for compile.Opinionated default: build the metric in pure Python with dspy.Predict(Assess) calls (transparent, version-controllable) and import LlamaIndex evaluators only for the RAG triad where the LlamaIndex evaluators are battle-tested. Avoid vendor SDKs whose rubrics you cannot inspect.
import dspy
class Assess(dspy.Signature):
"""One yes/no assessment."""
text = dspy.InputField()
question = dspy.InputField()
answer: bool = dspy.OutputField()
# Use a DIFFERENT family from the task model
judge_lm = dspy.LM("anthropic/claude-haiku-4") # task model is openai/gpt-4o
def make_metric(budget_words=150):
def metric(example, pred, trace=None):
with dspy.context(lm=judge_lm):
factual = dspy.Predict(Assess)(text=pred.answer,
question="Is every claim supported by the provided context?").answer
on_topic = dspy.Predict(Assess)(text=pred.answer,
question=f"Does this directly address: {example.question}?").answer
non_hedging = dspy.Predict(Assess)(text=pred.answer,
question="Does this commit to a position (no 'as an AI', no 'it depends')?").answer
n_words = len(pred.answer.split())
length_score = 1.0 if n_words <= budget_words else max(0.0, 1 - (n_words - budget_words)/budget_words)
if trace is not None:
return bool(factual and on_topic and non_hedging and n_words <= budget_words)
return (int(factual) + int(on_topic) + int(non_hedging)) / 3.0 * length_score
return metricThen: bias-probe → human-calibrate → save receipt → ship.
references/R1-source-evidence.md — DSPy / LlamaIndex / LangGraph source quotesreferences/R2-judge-bias-catalog.md — Judge bias taxonomy + mitigationsintermediate/operation_candidates.json — Operation registryCitations: [dspy.ai/learn/evaluation/metrics/], [dspy.ai/learn/optimization/overview/], [dspy.ai/cheatsheet/], [dspy.ai/api/optimizers/GEPA/overview/], [arxiv.org/abs/2507.19457], [arxiv.org/pdf/2506.02592], [arxiv.org/pdf/2509.26072], [developers.llamaindex.ai/python/framework-api-reference/evaluation/], [cookbook.openai.com/examples/evaluation/evaluate_rag_with_llamaindex].
© agentsope, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 4 other files (references) in skills/agentsop-metric-design of agentsope/SkillAlchemy.
Open the folder on GitHubat commit d0f0355
Agentsop Metric Design next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Agentsop Metric Design this skillagentsope/SkillAlchemy | 436 | — | ~6.4k | Automated safety check: Pass | MIT | |
| Advanced Evaluationguanyang/open-agent-hub | 977 | 2 repos | ~4.2k | Automated safety check: Pass | MIT | |
| Agentic Evalgithub/awesome-copilot | 40k | 3 repos | ~1.5k | Automated safety check: Pass | MIT | |
| Clawpathy AutoresearchClawBio/ClawBio | 1.2k | — | ~1.4k | Automated safety check: Pass | MIT | |
| Suede AI EvalJasonColapietro/suede-creator-skills | 127 | — | ~3.3k | Automated safety check: Pass | MIT | |
| Commerce Evalsanthropics/commerce-agents | 3.2k | — | ~1.8k | Automated safety check: Pass | Apache-2.0 |
guanyang/open-agent-hub
This skill should be used for advanced LLM evaluation: LLM-as-judge systems, direct scoring, pairwise comparison, rubric calibration, evaluator bias mitigation, confidence scoring, and automated…
github/awesome-copilot
Patterns and techniques for evaluating and improving AI agent outputs.
ClawBio/ClawBio
Eval-driven skill tuning. An agent skill from ClawBio/ClawBio.
JasonColapietro/suede-creator-skills
Suede AI eval design and coverage audit: AI-SPEC, failure-mode rubric with severity scoring, concrete pass/fail eval cases, coverage and infrastructure scores, and mechanical acceptance gates.
anthropics/commerce-agents
Authoring and running behavioral evals for a shopping or merchant agent, covering the case shape, authoring rules, code graders and judges, the run pattern, and poisoned fixtures.
aiskillstore/marketplace
This skill should be used when the user asks to "implement LLM-as-judge", "compare model outputs", "create evaluation rubrics", "mitigate evaluation bias", or mentions direct scoring, pairwise…
agentsope/SkillAlchemy
SOP for terminal-based, git-native AI pair programming with Aider (git work-tree + tree-sitter repo-map + edit-format + human-in-loop REPL).
agentsope/SkillAlchemy
Coder-agent working-file budget discipline: keep the editable working set (files you /add into writable context) under ~25k tokens, separate "read" from "edit", delegate breadth to a read-only…
agentsope/SkillAlchemy
Split a multi-call LM workflow by cognitive load, not by accuracy: let one strong model make the few reasoning decisions and a cheap model do the many mechanical executions (Aider architect+editor…
agentsope/SkillAlchemy
SOP for building multi-agent systems with CrewAI — role-based collaboration, sequential/hierarchical processes, Flows, memory, delegation.
agentsope/SkillAlchemy
SOP for building LLM applications on Dify — visual workflow + chatflow + agent + RAG knowledge base + plugin marketplace + observability, self-hostable.
agentsope/SkillAlchemy
Designs multiscale chunking for RAG by embedding small units for retrieval precision and returning larger context for synthesis.
Categories
Decomposed, multi-criteria metric design for LLM pipelines. An agent skill from agentsope/SkillAlchemy. Agentsop Metric Design is an agent skill from agentsope/SkillAlchemy. Decomposed, multi-criteria metric design for LLM pipelines.
Agentsop Metric Design fits situations like: tasks that involve LLM evaluation; tasks that involve Quizzes and assessments; tasks that involve Operations and SOPs.
Run `npx skills add agentsope/SkillAlchemy --skill agentsop-metric-design -a claude-code`. Or copy the skill folder (skills/agentsop-metric-design in agentsope/SkillAlchemy) into .claude/skills/agentsop-metric-design in your project. Claude Code loads it when a task matches its description.
Run `npx skills add agentsope/SkillAlchemy --skill agentsop-metric-design -a codex`. Or copy the skill folder (skills/agentsop-metric-design in agentsope/SkillAlchemy) into .agents/skills/agentsop-metric-design in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add agentsope/SkillAlchemy --skill agentsop-metric-design -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/agentsop-metric-design, .gemini/skills/agentsop-metric-design, .github/skills/agentsop-metric-design and .opencode/skills/agentsop-metric-design in your project.
SKILL.md names no scripts, command-line tools or credentials: Agentsop Metric Design is instructions for the agent only. Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Agentsop Metric Design is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 6.4k tokens (SKILL.md is roughly 26k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4.6k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Agentsop Metric Design: Advanced Evaluation (guanyang/open-agent-hub, 977 stars), Agentic Eval (github/awesome-copilot, 40k stars), Clawpathy Autoresearch (ClawBio/ClawBio, 1.2k stars) and Suede AI Eval (JasonColapietro/suede-creator-skills, 127 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
agentsope (a GitHub user) maintains it in agentsope/SkillAlchemy, which has 436 GitHub stars. The repository holds 46 skills in this directory. The repository was last updated on October 9, 2026.
Source: agentsope/SkillAlchemy on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.