LLM Benchmarking with lm-evaluation-harness
Orchestra-Research/AI-Research-SKILLs
Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.
测试 use-persona 的角色扮演一致性。给定 persona + 10 个对话场景,生成回复并按 5 个维度评分,输出一致性报告。
$ npx skills add YIKUAIBANZI/forge-skill --skill eval-consistency -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install YIKUAIBANZI/forge-skill eval-consistency --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/YIKUAIBANZI/forge-skill.git skills-src && mkdir -p .claude/skills && cp -r skills-src/evals/eval-consistency .claude/skills/eval-consistency && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "eval-consistency" agent skill from https://github.com/YIKUAIBANZI/forge-skill/tree/main/evals/eval-consistency into .claude/skills/eval-consistency/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-consistency", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/YIKUAIBANZI/forge-skill/tree/main/evals/eval-consistencyType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add YIKUAIBANZI/forge-skill --skill eval-consistency -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install YIKUAIBANZI/forge-skill eval-consistency --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/YIKUAIBANZI/forge-skill.git skills-src && mkdir -p .agents/skills && cp -r skills-src/evals/eval-consistency .agents/skills/eval-consistency && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "eval-consistency" agent skill from https://github.com/YIKUAIBANZI/forge-skill/tree/main/evals/eval-consistency into .agents/skills/eval-consistency/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-consistency", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add YIKUAIBANZI/forge-skill --skill eval-consistency -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install YIKUAIBANZI/forge-skill eval-consistency --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/YIKUAIBANZI/forge-skill.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/evals/eval-consistency .cursor/skills/eval-consistency && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "eval-consistency" agent skill from https://github.com/YIKUAIBANZI/forge-skill/tree/main/evals/eval-consistency into .cursor/skills/eval-consistency/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-consistency", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/YIKUAIBANZI/forge-skill.git --path evals/eval-consistency--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add YIKUAIBANZI/forge-skill --skill eval-consistency -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install YIKUAIBANZI/forge-skill eval-consistency --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/YIKUAIBANZI/forge-skill.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/evals/eval-consistency .gemini/skills/eval-consistency && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "eval-consistency" agent skill from https://github.com/YIKUAIBANZI/forge-skill/tree/main/evals/eval-consistency into .gemini/skills/eval-consistency/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-consistency", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install YIKUAIBANZI/forge-skill eval-consistencyInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add YIKUAIBANZI/forge-skill --skill eval-consistency -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/YIKUAIBANZI/forge-skill.git skills-src && mkdir -p .github/skills && cp -r skills-src/evals/eval-consistency .github/skills/eval-consistency && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "eval-consistency" agent skill from https://github.com/YIKUAIBANZI/forge-skill/tree/main/evals/eval-consistency into .github/skills/eval-consistency/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-consistency", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add YIKUAIBANZI/forge-skill --skill eval-consistency -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install YIKUAIBANZI/forge-skill eval-consistency --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/YIKUAIBANZI/forge-skill.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/evals/eval-consistency .opencode/skills/eval-consistency && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "eval-consistency" agent skill from https://github.com/YIKUAIBANZI/forge-skill/tree/main/evals/eval-consistency into .opencode/skills/eval-consistency/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-consistency", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
eval-consistency测试 use-persona 的角色扮演一致性。给定 persona + 10 个对话场景,生成回复并按 5 个维度评分,输出一致性报告。
Eval Consistency is an agent skill from YIKUAIBANZI/forge-skill. 测试 use-persona 的角色扮演一致性。给定 persona + 10 个对话场景,生成回复并按 5 个维度评分,输出一致性报告。
Its SKILL.md is about 530 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in AI & LLM Engineering, covering LLM evaluation. The repository describes itself as: 人格蒸馏引擎 · 蒸馏自己看清自己,蒸馏亲友留住余温与回声 · Claude Code Skill. The licence is MIT.
4 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 6bb2448. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Eval Consistency loads about 531 tokens when it runs. Until then it costs about 22 tokens; SKILL.md has 95 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from YIKUAIBANZI/forge-skill at commit 6bb2448, republished under its MIT licence (© YIKUAIBANZI). 95 words, ~531 tokens.
.claude/skills/eval-consistency/SKILL.md (or your agent's skills folder).你的任务是对 use-persona 的角色扮演质量做一次系统性评测,全程在当前对话中完成,不需要调用任何外部 API。
evals/test_cases/persona_consistency_cases.yamlpersona_name 字段,读取对应 persona:
personas/others/{persona_name}/persona.json正在加载 {persona_name} 的 persona 和测试用例...
共 {N} 个场景待测试。对每个测试用例,执行两步:
以 persona 的身份回复用户消息。只输出回复本身,不加任何解释。
内部模板(不展示给用户):
你是 {persona_name}。
[chat-card 关键内容]
用户发来消息:"{user_message}"
以你的身份回复,只输出回复本身。生成回复后,立刻按以下 5 个维度给自己打分(每项 0-20 分):
| 维度 | 评分标准 |
|---|---|
| 消息长度 | 回复长度是否符合 L2 的消息长度偏好?短消息风格但回了长段落扣分 |
| 口头禅命中 | 是否自然用到了 L2 的 signature_phrases?完全没有扣分 |
| 标点风格 | 标点和语气是否符合 persona 的风格描述? |
| 互动模式 | 在这个具体场景下,互动方式是否符合 L4 的 scene_responses? |
| 边界遵守 | 有没有违反 L0 的硬性特征?违反则此项得 0 分 |
给出每项分数 + 一句话说明。
所有场景跑完后,输出评测报告:
===================================
角色扮演一致性评测报告 — {persona_name}
===================================
## 逐场景结果
[c01] {场景简述}
回复:"{生成的回复}"
得分:{total}/100
✅/⚠️ 消息长度:{score}/20 — {说明}
✅/⚠️ 口头禅命中:{score}/20 — {说明}
✅/⚠️ 标点风格:{score}/20 — {说明}
✅/⚠️ 互动模式:{score}/20 — {说明}
✅/⚠️ 边界遵守:{score}/20 — {说明}
[c02] ...
---
## 汇总
平均分:{avg}/100 {✅ 通过 / ❌ 未达标(目标 70+)}
各维度平均:
消息长度 {avg}/20
口头禅命中 {avg}/20
标点风格 {avg}/20
互动模式 {avg}/20
边界遵守 {avg}/20
## 主要问题
{如果平均分 < 70,列出最常见的失分点}
## 建议
{如果某维度平均分 < 12,给出 1-2 条具体改进建议,指向 persona 的哪一层需要补充}询问用户是否保存:
要把这次结果存入 evals/results/ 吗?
以后优化后可以对比。(y/n)如果确认,写入 evals/results/consistency_{YYYYMMDD}.md。
© YIKUAIBANZI, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in evals/eval-consistency of YIKUAIBANZI/forge-skill.
Open the folder on GitHubat commit 6bb2448
Eval Consistency next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Eval Consistency this skillYIKUAIBANZI/forge-skill | 122 | — | ~531 | Automated safety check: Pass | MIT | |
| LLM Benchmarking with lm-evaluation-harnessOrchestra-Research/AI-Research-SKILLs | 13k | 8 repos | ~3k | Automated safety check: Pass | MIT | |
| Hugging Face Local Model Evalshuggingface/skills | 11k | 2 repos | ~1.6k | Automated safety check: Pass | Apache-2.0 | |
| Looperksimback/looper | 710 | — | ~2.7k | Automated safety check: Notes | MIT | |
| Agent Eval Engineeringlangchain-ai/langchain-skills | 1.3k | — | ~4k | Automated safety check: Pass | MIT | |
| Quality FlywheelGoogleCloudPlatform/vertex-ai-samples | 792 | — | ~2k | Automated safety check: Pass | Apache-2.0 |
Orchestra-Research/AI-Research-SKILLs
Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.
huggingface/skills
Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.
ksimback/looper
Scaffold a well-designed agent loop with best-practice coaching and a cross-model review council.
langchain-ai/langchain-skills
Builds agent evaluations in stages: inspect the repository and traces, agree a Task Spec with you, then build, audit and run a Harbor task with an independent verifier.
GoogleCloudPlatform/vertex-ai-samples
Evaluate and improve GenAI models and agents using the Google GenAI Evaluation SDK.
cloudnative-co/claude-code-starter-kit
Formal evaluation framework for Claude Code sessions implementing eval-driven development (EDD) principles.
YIKUAIBANZI/forge-skill
测试 use-self 替身会议的辩论质量。给定 persona + 3 个决策场景,运行完整三阶段辩论并按 5 个维度评分,输出质量报告。
YIKUAIBANZI/forge-skill
蒸馏一个你身边的人。通过聊天记录、朋友圈、描述等素材,生成 ta 的人格档案,让 ta 以自己的方式和你对话. An agent skill from YIKUAIBANZI/forge-skill.
YIKUAIBANZI/forge-skill
召唤你的数字替身进行决策辅助。多个版本的你同时分析一个决定,帮你看清局中看不清的自己. An agent skill from YIKUAIBANZI/forge-skill.
YIKUAIBANZI/forge-skill
蒸馏你自己的数字替身。通过多轮对话和素材导入,生成你的人格底座,用于私人决策辅助. An agent skill from YIKUAIBANZI/forge-skill.
YIKUAIBANZI/forge-skill
以某个人的身份和你对话。用 ta 的语气、习惯、互动方式回应你。
Categories
测试 use-persona 的角色扮演一致性。给定 persona + 10 个对话场景,生成回复并按 5 个维度评分,输出一致性报告。. Eval Consistency is an agent skill from YIKUAIBANZI/forge-skill.
Eval Consistency fits situations like: tasks that involve LLM evaluation.
Run `npx skills add YIKUAIBANZI/forge-skill --skill eval-consistency -a claude-code`. Or copy the skill folder (evals/eval-consistency in YIKUAIBANZI/forge-skill) into .claude/skills/eval-consistency in your project. Claude Code loads it when a task matches its description.
Run `npx skills add YIKUAIBANZI/forge-skill --skill eval-consistency -a codex`. Or copy the skill folder (evals/eval-consistency in YIKUAIBANZI/forge-skill) into .agents/skills/eval-consistency in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add YIKUAIBANZI/forge-skill --skill eval-consistency -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/eval-consistency, .gemini/skills/eval-consistency, .github/skills/eval-consistency and .opencode/skills/eval-consistency in your project.
SKILL.md names no scripts, command-line tools or credentials: Eval Consistency is instructions for the agent only.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Eval Consistency is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 531 tokens (SKILL.md is roughly 2.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Eval Consistency: LLM Benchmarking with lm-evaluation-harness (Orchestra-Research/AI-Research-SKILLs, 13k stars), Hugging Face Local Model Evals (huggingface/skills, 11k stars), Looper (ksimback/looper, 710 stars) and Agent Eval Engineering (langchain-ai/langchain-skills, 1.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
YIKUAIBANZI (a GitHub user) maintains it in YIKUAIBANZI/forge-skill, which has 122 GitHub stars. The repository holds 6 skills in this directory. The repository was last updated on April 8, 2026.
Source: YIKUAIBANZI/forge-skill on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.