Agent skill

Eval Consistency

by YIKUAIBANZI in YIKUAIBANZI/forge-skill

测试 use-persona 的角色扮演一致性。给定 persona + 10 个对话场景,生成回复并按 5 个维度评分,输出一致性报告。

MITAuto-check passedAI & LLM Engineering

Install Eval Consistency

skills CLI
$ npx skills add YIKUAIBANZI/forge-skill --skill eval-consistency -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install YIKUAIBANZI/forge-skill eval-consistency --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/YIKUAIBANZI/forge-skill.git skills-src && mkdir -p .claude/skills && cp -r skills-src/evals/eval-consistency .claude/skills/eval-consistency && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
eval-consistency
GitHub stars
122
Token cost
~531 tokens
SKILL.md length
95 words
Files
1
Skills in repo
6
Repo updated
First seen
Licence
MIT

At a glance

测试 use-persona 的角色扮演一致性。给定 persona + 10 个对话场景,生成回复并按 5 个维度评分,输出一致性报告。

  • Works in 4 steps: :加载测试资源 → :逐场景测试 → :输出完整报告 → …
  • Tasks that involve LLM evaluation
  • SKILL.md covers Step 0:加载测试资源, Step 1:逐场景测试, Step 2:输出完整报告 and Step 3:保存结果(可选), plus 1 more section
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Eval Consistency is an agent skill from YIKUAIBANZI/forge-skill. 测试 use-persona 的角色扮演一致性。给定 persona + 10 个对话场景,生成回复并按 5 个维度评分,输出一致性报告。

Its SKILL.md is about 530 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering LLM evaluation. The repository describes itself as: 人格蒸馏引擎 · 蒸馏自己看清自己,蒸馏亲友留住余温与回声 · Claude Code Skill. The licence is MIT.

When your agent uses it

  • Tasks that involve LLM evaluation

Example prompts

  • “/eval-consistency”

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. :加载测试资源
  2. :逐场景测试
  3. :输出完整报告
  4. :保存结果(可选)

What it can do on your machine

Read from SKILL.md and the folder at commit 6bb2448. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Eval Consistency loads about 531 tokens when it runs. Until then it costs about 22 tokens; SKILL.md has 95 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~22
When it runs · the whole SKILL.md, loaded when a task matches
~531

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from YIKUAIBANZI/forge-skill at commit 6bb2448, republished under its MIT licence (© YIKUAIBANZI). 95 words, ~531 tokens.

Download SKILL.mdSave it as .claude/skills/eval-consistency/SKILL.md (or your agent's skills folder).
name
eval-consistency
description
测试 use-persona 的角色扮演一致性。给定 persona + 10 个对话场景,生成回复并按 5 个维度评分,输出一致性报告。
trigger
当用户说"/eval-consistency"、"测试角色扮演一致性"、"一致性评测"时触发
tools
Read, Write, Glob

/eval-consistency — 角色扮演一致性评测

你的任务是对 use-persona 的角色扮演质量做一次系统性评测,全程在当前对话中完成,不需要调用任何外部 API。


Step 0:加载测试资源

  1. 读取测试用例文件:evals/test_cases/persona_consistency_cases.yaml
  2. 根据 persona_name 字段,读取对应 persona: personas/others/{persona_name}/persona.json
  3. 从 persona.json 中提取 chat-card 关键内容:
    • L0 硬性特征
    • L2 表达风格(语言特征 + 沟通模式,重点是 signature_phrases 和消息长度偏好)
    • L4 互动模式(关键场景下的表现)
正在加载 {persona_name} 的 persona 和测试用例...
共 {N} 个场景待测试。

Step 1:逐场景测试

对每个测试用例,执行两步:

1a. 生成角色扮演回复

以 persona 的身份回复用户消息。只输出回复本身,不加任何解释。

内部模板(不展示给用户):

你是 {persona_name}。
[chat-card 关键内容]

用户发来消息:"{user_message}"

以你的身份回复,只输出回复本身。
1b. 评分(内部执行,立即给出)

生成回复后,立刻按以下 5 个维度给自己打分(每项 0-20 分):

维度评分标准
消息长度回复长度是否符合 L2 的消息长度偏好?短消息风格但回了长段落扣分
口头禅命中是否自然用到了 L2 的 signature_phrases?完全没有扣分
标点风格标点和语气是否符合 persona 的风格描述?
互动模式在这个具体场景下,互动方式是否符合 L4 的 scene_responses?
边界遵守有没有违反 L0 的硬性特征?违反则此项得 0 分

给出每项分数 + 一句话说明。


Step 2:输出完整报告

所有场景跑完后,输出评测报告:

===================================
角色扮演一致性评测报告 — {persona_name}
===================================

## 逐场景结果

[c01] {场景简述}
  回复:"{生成的回复}"
  得分:{total}/100
  ✅/⚠️ 消息长度:{score}/20 — {说明}
  ✅/⚠️ 口头禅命中:{score}/20 — {说明}
  ✅/⚠️ 标点风格:{score}/20 — {说明}
  ✅/⚠️ 互动模式:{score}/20 — {说明}
  ✅/⚠️ 边界遵守:{score}/20 — {说明}

[c02] ...

---

## 汇总

平均分:{avg}/100  {✅ 通过 / ❌ 未达标(目标 70+)}

各维度平均:
  消息长度    {avg}/20
  口头禅命中  {avg}/20
  标点风格    {avg}/20
  互动模式    {avg}/20
  边界遵守    {avg}/20

## 主要问题
{如果平均分 < 70,列出最常见的失分点}

## 建议
{如果某维度平均分 < 12,给出 1-2 条具体改进建议,指向 persona 的哪一层需要补充}

Step 3:保存结果(可选)

询问用户是否保存:

要把这次结果存入 evals/results/ 吗?
以后优化后可以对比。(y/n)

如果确认,写入 evals/results/consistency_{YYYYMMDD}.md。


注意

  • 全程不需要 API Key:评分是你自己执行的,不是另起一个 LLM
  • 评分要诚实:对自己生成的回复该扣分就扣分,不要因为是自己生成的就打高分
  • 用例是基于小美的,如果用户指定了其他 persona,根据那个 persona 的 L2/L4 调整评分标准

© YIKUAIBANZI, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in evals/eval-consistency of YIKUAIBANZI/forge-skill.

Open the folder on GitHubat commit 6bb2448

Compare with similar skills

Eval Consistency next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Eval Consistency compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Eval Consistency this skillYIKUAIBANZI/forge-skill122—~531Automated safety check: PassMIT
LLM Benchmarking with lm-evaluation-harnessOrchestra-Research/AI-Research-SKILLs13k8 repos~3kAutomated safety check: PassMIT
Hugging Face Local Model Evalshuggingface/skills11k2 repos~1.6kAutomated safety check: PassApache-2.0
Looperksimback/looper710—~2.7kAutomated safety check: NotesMIT
Agent Eval Engineeringlangchain-ai/langchain-skills1.3k—~4kAutomated safety check: PassMIT
Quality FlywheelGoogleCloudPlatform/vertex-ai-samples792—~2kAutomated safety check: PassApache-2.0

Similar skills

  • LLM Benchmarking with lm-evaluation-harness

    Orchestra-Research/AI-Research-SKILLs

    Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.

    13k GitHub starsUsed in 8 repos~3k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.

    11k GitHub starsUsed in 2 repos~1.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Looper

    ksimback/looper

    Scaffold a well-designed agent loop with best-practice coaching and a cross-model review council.

    710 GitHub stars~2.7k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check: notes
  • Agent Eval Engineering

    langchain-ai/langchain-skills

    Official

    Builds agent evaluations in stages: inspect the repository and traces, agree a Task Spec with you, then build, audit and run a Harbor task with an independent verifier.

    1.3k GitHub stars~4k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Quality Flywheel

    GoogleCloudPlatform/vertex-ai-samples

    Evaluate and improve GenAI models and agents using the Google GenAI Evaluation SDK.

    792 GitHub stars~2k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Eval Harness

    cloudnative-co/claude-code-starter-kit

    Formal evaluation framework for Claude Code sessions implementing eval-driven development (EDD) principles.

    153 GitHub starsUsed in 9 repos~1.3k tokens
    AI & LLM EngineeringAuto-check passed

More from YIKUAIBANZI/forge-skill

  • Eval Debate

    YIKUAIBANZI/forge-skill

    测试 use-self 替身会议的辩论质量。给定 persona + 3 个决策场景,运行完整三阶段辩论并按 5 个维度评分,输出质量报告。

    122 GitHub stars~740 tokensUpdated 6 mo ago
    Auto-check passed
  • Forge Persona

    YIKUAIBANZI/forge-skill

    蒸馏一个你身边的人。通过聊天记录、朋友圈、描述等素材,生成 ta 的人格档案,让 ta 以自己的方式和你对话. An agent skill from YIKUAIBANZI/forge-skill.

    122 GitHub stars~543 tokensUpdated 6 mo ago
    Auto-check passed
  • Use Self

    YIKUAIBANZI/forge-skill

    召唤你的数字替身进行决策辅助。多个版本的你同时分析一个决定,帮你看清局中看不清的自己. An agent skill from YIKUAIBANZI/forge-skill.

    122 GitHub stars~593 tokensUpdated 6 mo ago
    Auto-check passed
  • Forge Self

    YIKUAIBANZI/forge-skill

    蒸馏你自己的数字替身。通过多轮对话和素材导入,生成你的人格底座,用于私人决策辅助. An agent skill from YIKUAIBANZI/forge-skill.

    122 GitHub stars~449 tokensUpdated 6 mo ago
    Auto-check passed
  • Use Persona

    YIKUAIBANZI/forge-skill

    以某个人的身份和你对话。用 ta 的语气、习惯、互动方式回应你。

    122 GitHub stars~337 tokensUpdated 6 mo ago
    Auto-check passed

Questions about Eval Consistency

What does Eval Consistency do?

测试 use-persona 的角色扮演一致性。给定 persona + 10 个对话场景,生成回复并按 5 个维度评分,输出一致性报告。. Eval Consistency is an agent skill from YIKUAIBANZI/forge-skill.

When should I use Eval Consistency?

Eval Consistency fits situations like: tasks that involve LLM evaluation.

How do I install Eval Consistency in Claude Code?

Run `npx skills add YIKUAIBANZI/forge-skill --skill eval-consistency -a claude-code`. Or copy the skill folder (evals/eval-consistency in YIKUAIBANZI/forge-skill) into .claude/skills/eval-consistency in your project. Claude Code loads it when a task matches its description.

How do I install Eval Consistency in Codex?

Run `npx skills add YIKUAIBANZI/forge-skill --skill eval-consistency -a codex`. Or copy the skill folder (evals/eval-consistency in YIKUAIBANZI/forge-skill) into .agents/skills/eval-consistency in your project. Codex loads it when a task matches its description.

Can I use Eval Consistency in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add YIKUAIBANZI/forge-skill --skill eval-consistency -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/eval-consistency, .gemini/skills/eval-consistency, .github/skills/eval-consistency and .opencode/skills/eval-consistency in your project.

What does Eval Consistency need to run?

SKILL.md names no scripts, command-line tools or credentials: Eval Consistency is instructions for the agent only.

Does Eval Consistency access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Eval Consistency safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Eval Consistency use?

Eval Consistency is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Eval Consistency use?

About 531 tokens (SKILL.md is roughly 2.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Eval Consistency?

Skills that share tags, products or a category with Eval Consistency: LLM Benchmarking with lm-evaluation-harness (Orchestra-Research/AI-Research-SKILLs, 13k stars), Hugging Face Local Model Evals (huggingface/skills, 11k stars), Looper (ksimback/looper, 710 stars) and Agent Eval Engineering (langchain-ai/langchain-skills, 1.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Eval Consistency?

YIKUAIBANZI (a GitHub user) maintains it in YIKUAIBANZI/forge-skill, which has 122 GitHub stars. The repository holds 6 skills in this directory. The repository was last updated on April 8, 2026.

Source: YIKUAIBANZI/forge-skill on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.