Agent skill

Eval Debate

by YIKUAIBANZI in YIKUAIBANZI/forge-skill

测试 use-self 替身会议的辩论质量。给定 persona + 3 个决策场景,运行完整三阶段辩论并按 5 个维度评分,输出质量报告。

MITAuto-check passedAI & LLM Engineering

Install Eval Debate

skills CLI
$ npx skills add YIKUAIBANZI/forge-skill --skill eval-debate -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install YIKUAIBANZI/forge-skill eval-debate --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/YIKUAIBANZI/forge-skill.git skills-src && mkdir -p .claude/skills && cp -r skills-src/evals/eval-debate .claude/skills/eval-debate && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
eval-debate
GitHub stars
122
Token cost
~740 tokens
SKILL.md length
161 words
Files
1
Skills in repo
6
Repo updated
First seen
Licence
MIT

At a glance

测试 use-self 替身会议的辩论质量。给定 persona + 3 个决策场景,运行完整三阶段辩论并按 5 个维度评分,输出质量报告。

  • Works in 5 steps: :加载测试资源 → :逐场景运行辩论 → :评分 → …
  • Tasks that involve LLM evaluation
  • SKILL.md covers Step 0:加载测试资源, Step 1:逐场景运行辩论, Step 2:评分 and Step 3:输出报告, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Eval Debate is an agent skill from YIKUAIBANZI/forge-skill. 测试 use-self 替身会议的辩论质量。给定 persona + 3 个决策场景,运行完整三阶段辩论并按 5 个维度评分,输出质量报告。

Its SKILL.md is about 740 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering LLM evaluation. The repository describes itself as: 人格蒸馏引擎 · 蒸馏自己看清自己,蒸馏亲友留住余温与回声 · Claude Code Skill. The licence is MIT.

When your agent uses it

  • Tasks that involve LLM evaluation

Example prompts

  • “/eval-debate”

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. :加载测试资源
  2. :逐场景运行辩论
  3. :评分
  4. :输出报告
  5. :保存结果(可选)

What it can do on your machine

Read from SKILL.md and the folder at commit 6bb2448. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Eval Debate loads about 740 tokens when it runs. Until then it costs about 21 tokens; SKILL.md has 161 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~21
When it runs · the whole SKILL.md, loaded when a task matches
~740

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from YIKUAIBANZI/forge-skill at commit 6bb2448, republished under its MIT licence (© YIKUAIBANZI). 161 words, ~740 tokens.

Download SKILL.mdSave it as .claude/skills/eval-debate/SKILL.md (or your agent's skills folder).
name
eval-debate
description
测试 use-self 替身会议的辩论质量。给定 persona + 3 个决策场景,运行完整三阶段辩论并按 5 个维度评分,输出质量报告。
trigger
当用户说"/eval-debate"、"测试辩论质量"、"辩论评测"时触发
tools
Read, Write, Glob

/eval-debate — 替身会议辩论质量评测

你的任务是对 use-self 替身会议的输出质量做一次系统性评测,全程在当前对话中完成,不需要调用任何外部 API。


Step 0:加载测试资源

  1. 读取测试用例文件:evals/test_cases/debate_quality_cases.yaml
  2. 根据 persona_name 字段,读取对应 persona: personas/self/{persona_name}/persona.json
  3. 提取 decision-card 关键内容:
    • L0 底线(bottom_line)
    • L2 语言风格(language_style + signature_phrases)
    • L3 决策参数(8 个维度的分值)
    • L4 价值观与盲区(blind_spots + emotional_triggers)
正在加载 {persona_name} 的 persona 和测试用例...
共 {N} 个决策场景待测试。

Step 1:逐场景运行辩论

对每个测试用例,执行完整三阶段流程:

Phase 1:并行独立分析(3 个变体)

基于 decision-card 中的 L3 参数,生成 3 个变体并各自独立分析:

变体设置(固定,评测用):

  • 🔵 稳健的你:risk_appetite -3,action_bias -2,loss_aversion +2
  • 🟢 果断的你:risk_appetite +3,action_bias +3,information_need -2
  • 🔴 长线的你:time_horizon +4,loss_aversion -2,action_bias +1

每个变体按 use-self/prompts/phase1_independent.md 的格式输出:

  • 【我的判断】:明确表态,不能含糊
  • 【为什么】:≤3 个具体理由
  • 【我最担心的是】:具体情境
  • 【我最期待的是】:具体情境
  • 【我想问自己】:一个核心问题

信息隔离:每个变体只能看到自己的参数偏移,不知道其他变体说了什么。

Phase 2:质询

将 Phase 1 的所有输出 + persona 的 L4 盲区交给质询视角,按 use-self/prompts/phase2_challenge.md 执行:

  • 对每个变体找出最尖锐的质疑(隐含假设/幻觉/回避)
  • 识别跨变体矛盾
  • 用 L4 盲区做最后一问
Phase 3:综合

按 use-self/prompts/phase3_synthesis.md 生成综合报告,使用用户的 L2 语言风格。


Step 2:评分

每个场景跑完后,立刻按 5 个维度评分(每项 0-20 分):

维度评分标准
变体区分度Phase 1 的 3 个变体立场是否有实质性差异?都说"两边各有道理"= 0 分;立场明确对立且理由具体 = 满分
质询深度Phase 2 是否指出了具体假设和盲区?"你没考虑到..." = 低分;"你说的 X 假设了 Y,但 Y 不成立,因为 Z" = 高分
参数一致性各变体的发言是否与偏移后的参数一致?稳健变体的发言是否明显更保守?
综合覆盖度Phase 3 是否有代价清单?是否提出了具体的待搞清楚的问题?还是只是 Phase 1 的复述?
用户语言风格所有输出语气是否符合 persona 的 L2?出现"综上所述"、"建议您"等顾问句式扣分

参照测试用例的 expected_variant_stances 和 evaluation_criteria 给分。


Step 3:输出报告

所有场景跑完后,输出评测报告:

===================================
替身会议辩论质量评测报告 — {persona_name}
===================================

## 逐场景结果

### [d01] {场景标题}

**Phase 1 摘要:**
- 🔵 稳健的你:{判断一句话}
- 🟢 果断的你:{判断一句话}
- 🔴 长线的你:{判断一句话}

**Phase 2 质询摘要:**
{最有价值的一条质疑}

**Phase 3 综合摘要:**
{代价清单里最关键的一条}

**评分:{total}/100**
  ✅/⚠️ 变体区分度:{score}/20 — {说明}
  ✅/⚠️ 质询深度:{score}/20 — {说明}
  ✅/⚠️ 参数一致性:{score}/20 — {说明}
  ✅/⚠️ 综合覆盖度:{score}/20 — {说明}
  ✅/⚠️ 用户语言风格:{score}/20 — {说明}

### [d02] ...
### [d03] ...

---

## 汇总

平均分:{avg}/100

各维度平均:
  变体区分度    {avg}/20
  质询深度      {avg}/20
  参数一致性    {avg}/20
  综合覆盖度    {avg}/20
  用户语言风格  {avg}/20

## 主要问题
{失分最多的维度 + 具体表现}

## 建议
{针对失分维度的改进方向,指向哪个 prompt 文件或 persona 层级需要调整}

Step 4:保存结果(可选)

要把这次结果存入 evals/results/ 吗?(y/n)

如果确认,写入 evals/results/debate_{YYYYMMDD}.md。


注意

  • 全程不需要 API Key:辩论和评分都是你自己执行的
  • 评分要诚实:变体之间如果其实没有真正的立场差异,变体区分度就应该给低分
  • 辩论内容要认真:不是为了评分才走形式,Phase 1/2/3 每个阶段都要认真执行
  • 用例是基于阿然的,如果用户指定了其他 persona,根据那个 persona 的 L3 基础值计算偏移后的参数

© YIKUAIBANZI, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in evals/eval-debate of YIKUAIBANZI/forge-skill.

Open the folder on GitHubat commit 6bb2448

Compare with similar skills

Eval Debate next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Eval Debate compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Eval Debate this skillYIKUAIBANZI/forge-skill122—~740Automated safety check: PassMIT
LLM Benchmarking with lm-evaluation-harnessOrchestra-Research/AI-Research-SKILLs13k8 repos~3kAutomated safety check: PassMIT
Hugging Face Local Model Evalshuggingface/skills11k2 repos~1.6kAutomated safety check: PassApache-2.0
Looperksimback/looper710—~2.7kAutomated safety check: NotesMIT
Agent Eval Engineeringlangchain-ai/langchain-skills1.3k—~4kAutomated safety check: PassMIT
Quality FlywheelGoogleCloudPlatform/vertex-ai-samples792—~2kAutomated safety check: PassApache-2.0

Similar skills

  • LLM Benchmarking with lm-evaluation-harness

    Orchestra-Research/AI-Research-SKILLs

    Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.

    13k GitHub starsUsed in 8 repos~3k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.

    11k GitHub starsUsed in 2 repos~1.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Looper

    ksimback/looper

    Scaffold a well-designed agent loop with best-practice coaching and a cross-model review council.

    710 GitHub stars~2.7k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check: notes
  • Agent Eval Engineering

    langchain-ai/langchain-skills

    Official

    Builds agent evaluations in stages: inspect the repository and traces, agree a Task Spec with you, then build, audit and run a Harbor task with an independent verifier.

    1.3k GitHub stars~4k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Quality Flywheel

    GoogleCloudPlatform/vertex-ai-samples

    Evaluate and improve GenAI models and agents using the Google GenAI Evaluation SDK.

    792 GitHub stars~2k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Eval Harness

    cloudnative-co/claude-code-starter-kit

    Formal evaluation framework for Claude Code sessions implementing eval-driven development (EDD) principles.

    153 GitHub starsUsed in 9 repos~1.3k tokens
    AI & LLM EngineeringAuto-check passed

More from YIKUAIBANZI/forge-skill

  • Eval Consistency

    YIKUAIBANZI/forge-skill

    测试 use-persona 的角色扮演一致性。给定 persona + 10 个对话场景,生成回复并按 5 个维度评分,输出一致性报告。

    122 GitHub stars~531 tokensUpdated 6 mo ago
    Auto-check passed
  • Forge Persona

    YIKUAIBANZI/forge-skill

    蒸馏一个你身边的人。通过聊天记录、朋友圈、描述等素材,生成 ta 的人格档案,让 ta 以自己的方式和你对话. An agent skill from YIKUAIBANZI/forge-skill.

    122 GitHub stars~543 tokensUpdated 6 mo ago
    Auto-check passed
  • Use Self

    YIKUAIBANZI/forge-skill

    召唤你的数字替身进行决策辅助。多个版本的你同时分析一个决定,帮你看清局中看不清的自己. An agent skill from YIKUAIBANZI/forge-skill.

    122 GitHub stars~593 tokensUpdated 6 mo ago
    Auto-check passed
  • Forge Self

    YIKUAIBANZI/forge-skill

    蒸馏你自己的数字替身。通过多轮对话和素材导入,生成你的人格底座,用于私人决策辅助. An agent skill from YIKUAIBANZI/forge-skill.

    122 GitHub stars~449 tokensUpdated 6 mo ago
    Auto-check passed
  • Use Persona

    YIKUAIBANZI/forge-skill

    以某个人的身份和你对话。用 ta 的语气、习惯、互动方式回应你。

    122 GitHub stars~337 tokensUpdated 6 mo ago
    Auto-check passed

Questions about Eval Debate

What does Eval Debate do?

测试 use-self 替身会议的辩论质量。给定 persona + 3 个决策场景,运行完整三阶段辩论并按 5 个维度评分,输出质量报告。. Eval Debate is an agent skill from YIKUAIBANZI/forge-skill.

When should I use Eval Debate?

Eval Debate fits situations like: tasks that involve LLM evaluation.

How do I install Eval Debate in Claude Code?

Run `npx skills add YIKUAIBANZI/forge-skill --skill eval-debate -a claude-code`. Or copy the skill folder (evals/eval-debate in YIKUAIBANZI/forge-skill) into .claude/skills/eval-debate in your project. Claude Code loads it when a task matches its description.

How do I install Eval Debate in Codex?

Run `npx skills add YIKUAIBANZI/forge-skill --skill eval-debate -a codex`. Or copy the skill folder (evals/eval-debate in YIKUAIBANZI/forge-skill) into .agents/skills/eval-debate in your project. Codex loads it when a task matches its description.

Can I use Eval Debate in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add YIKUAIBANZI/forge-skill --skill eval-debate -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/eval-debate, .gemini/skills/eval-debate, .github/skills/eval-debate and .opencode/skills/eval-debate in your project.

What does Eval Debate need to run?

SKILL.md names no scripts, command-line tools or credentials: Eval Debate is instructions for the agent only.

Does Eval Debate access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Eval Debate safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Eval Debate use?

Eval Debate is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Eval Debate use?

About 740 tokens (SKILL.md is roughly 3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Eval Debate?

Skills that share tags, products or a category with Eval Debate: LLM Benchmarking with lm-evaluation-harness (Orchestra-Research/AI-Research-SKILLs, 13k stars), Hugging Face Local Model Evals (huggingface/skills, 11k stars), Looper (ksimback/looper, 710 stars) and Agent Eval Engineering (langchain-ai/langchain-skills, 1.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Eval Debate?

YIKUAIBANZI (a GitHub user) maintains it in YIKUAIBANZI/forge-skill, which has 122 GitHub stars. The repository holds 6 skills in this directory. The repository was last updated on April 8, 2026.

Source: YIKUAIBANZI/forge-skill on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.