Open-Science Skill Creator
aipoch/open-science
Creates, revises, evaluates and publishes skills in the Open-Science app through its native host.skills composer, with optional test prompts and benchmarks.
Scores SKILL.md files on a nine-dimension rubric, then improves them in a keep-or-revert loop with independent judge agents, test prompts, git history and human checkpoints.
SKILL.md written in Chinese; this summary is our English description.
$ npx skills add alchaincyf/darwin-skill --skill darwin-skill -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install alchaincyf/darwin-skill darwin-skill --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
Claude Code skills documentation · loads skills from .claude/skills/
Install the "darwin-skill" agent skill from https://github.com/alchaincyf/darwin-skill/tree/master into .claude/skills/darwin-skill/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "darwin-skill", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add alchaincyf/darwin-skill --skill darwin-skill -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install alchaincyf/darwin-skill darwin-skill --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "darwin-skill" agent skill from https://github.com/alchaincyf/darwin-skill/tree/master into .agents/skills/darwin-skill/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "darwin-skill", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add alchaincyf/darwin-skill --skill darwin-skill -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install alchaincyf/darwin-skill darwin-skill --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "darwin-skill" agent skill from https://github.com/alchaincyf/darwin-skill/tree/master into .cursor/skills/darwin-skill/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "darwin-skill", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add alchaincyf/darwin-skill --skill darwin-skill -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install alchaincyf/darwin-skill darwin-skill --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "darwin-skill" agent skill from https://github.com/alchaincyf/darwin-skill/tree/master into .gemini/skills/darwin-skill/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "darwin-skill", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install alchaincyf/darwin-skill darwin-skillInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add alchaincyf/darwin-skill --skill darwin-skill -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "darwin-skill" agent skill from https://github.com/alchaincyf/darwin-skill/tree/master into .github/skills/darwin-skill/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "darwin-skill", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add alchaincyf/darwin-skill --skill darwin-skill -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install alchaincyf/darwin-skill darwin-skill --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "darwin-skill" agent skill from https://github.com/alchaincyf/darwin-skill/tree/master into .opencode/skills/darwin-skill/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "darwin-skill", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
darwin-skillScores SKILL.md files on a nine-dimension rubric, then improves them in a keep-or-revert loop with independent judge agents, test prompts, git history and human checkpoints.
The agent takes one SKILL.md at a time and scores it on nine dimensions worth 100 points in total: structure (frontmatter, workflow clarity, failure modes, checkpoints, actionable specificity, resources), effectiveness (overall architecture and measured performance on a few test prompts) and a meta-skill dimension that rewards an explicit list of things not to do. Absolute scores are used only to decide which skill is weakest and should go first.
Changes are kept or reverted by comparison, not by raw score: judge noise of about 8 points was observed, so each edit is compared against the original by the same judge, with an odd number of judges deciding by majority. Separate judge agents score blindly so the agent is not grading its own edit, git tracks every version, the loop stops on its own when gains diminish, and the agent pauses for your confirmation after each skill. It ends by producing a visual result card.
The design borrows from Karpathy's autoresearch idea of a single editable asset, a ratchet that keeps only improvements and independent scoring, and cites two Microsoft Research papers, SkillLens and SkillOpt, for the rubric and the validation checks. The skill text is written in Chinese.
7 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 8a8b662. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/, which the agent can run.
Shell commands in SKILL.md call:
gitclaudecursorcodexnodenpxpipFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
arxiv.orggithub.commicrosoft.github.ioFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Darwin Skill Optimizer loads about 4.7k tokens when it runs, and up to ~7.6k if it reads all its reference files. Until then it costs about 182 tokens; SKILL.md has 1,008 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from alchaincyf/darwin-skill at commit 8a8b662, republished under its MIT licence (© alchaincyf). 1,008 words, ~4,665 tokens.
.claude/skills/darwin-skill/SKILL.md (or your agent's skills folder). This skill also uses 41 other files; get the full folder from GitHub.v2.1 · 2026-06-10 — keep/revert 棘轮从「绝对分数 delta」改为「paired 同-judge 比较 + 奇数 N 多数决」(绝对分数 ±8 judge 噪音淹没保守编辑的真实增益、是 false-revert 源;within-judge 比较消除换尺污染)。绝对分数降级为 triage-only。 v2.0 · 2026-05-28 — 吸收 Microsoft Research SkillLens(arXiv 2605.23899)的 9 维评分药方 + SkillOpt(arXiv 2605.23904)的 validation-gated 验证机制 + human in the loop 三层守关。
借鉴 Karpathy autoresearch 的自主实验循环,对 skills 进行持续优化。 核心理念:评估 → 改进 → 实测验证 → 人类确认 → 保留或回滚 → 生成成果卡片 GitHub: https://github.com/alchaincyf/darwin-skill
autoresearch 的精髓:
与纯结构审查的区别:不只看 SKILL.md 写得规不规范,更看改完后实际跑出来的效果是否更好。
设计依据:基于 SkillLens 论文(arXiv 2605.23899)实证发现——LLM-as-judge 评估 skill 质量准确率仅 46.4%(接近随机),加入 meta-skill 三维度后提升到 73.8%。本 rubric 强化 dim3 / dim5 评分标准,新增 dim9「反例与黑名单」,权重平衡到 100。目的:让评分对真实质量更敏感,减少 LLM judge 的乐观偏差。
| # | 维度 | 权重 | 评分标准 |
|---|---|---|---|
| 1 | Frontmatter质量 | 7 | name规范、description包含做什么+何时用+触发词、≤1024字符、禁结尾加"灵活应用/根据情况判断"等空话尾巴 |
| 2 | 工作流清晰度 | 12 | 步骤明确可执行、有序号、每步有明确输入/输出 |
| 3 | 失败模式编码 | 12 | 必须显式编码失败模式(写出"如果 X 失败 → Y"的明确分支);有fallback路径、错误恢复;只写正向流程而不写失败分支扣 ≥3 分(SkillLens meta-skill 维度) |
| 4 | 检查点设计 | 6 | 关键决策前有用户确认、防止自主失控;检查点必须显性标记(🔴/STOP/CHECKPOINT),仅靠"如果...建议..."措辞不算 |
| 5 | 可执行具体性 | 18 | 不模糊、有具体参数/格式/示例、可直接执行;禁止"建议/可以考虑/根据情况/灵活把握/视情况而定"等软化措辞——出现 ≥3 处扣 ≥3 分(SkillLens actionable specificity 维度) |
| 6 | 资源整合度 | 4 | references/scripts/assets引用正确、路径可达 |
| # | 维度 | 权重 | 评分标准 |
|---|---|---|---|
| 7 | 整体架构 | 12 | 结构层次清晰、不冗余不遗漏、与花叔生态一致;冗余/AI腔废话段落(说白了/换句话说/首先其次综上等花叔禁用词)出现一处扣 1 分 |
| 8 | 实测表现 | 23 | 用测试prompt跑一遍,输出质量是否符合skill宣称的能力 |
| # | 维度 | 权重 | 评分标准 |
|---|---|---|---|
| 9 | 反例与黑名单 | 6 | skill 必须有"不要做什么"的反例清单;只写"应该做 X"没有"不要做 Y"扣 ≥3 分;红灯/危险动作/反模式应单独章节列出(SkillLens risk-action blacklist 维度) |
rubric 设计依据来自 SkillLens 论文(arXiv 2605.23899) + 本机 controlled study:
结论:rubric 能识别 gross degradation,但 fine-grained quality difference 仍不可信,重要决策必须人审。
→ 详细论文证据 + 5 judges 完整数据 + HL 实战案例数字见 references/skilllens-evidence.md
这是与纯结构评分最大的区别。评分方式:
若子 agent 不可用(超时/资源限制),退化为「干跑验证」:读完 skill 后模拟一个典型 prompt 的执行思路,判断流程是否合理;必须在 results.tsv 标注 dry_run。dry_run 比例 > 30% → 评估失效警告(来自本机 controlled study:dim8 实测维度权重 23%,无 full_test 验证时分数不可信)。
skill 应当能在 Claude Code / Codex / Cursor / OpenClaw / Hermes / Gemini CLI / OpenCode 等 50+ skills-compatible runtime 通用——否则其他 agent 解析时会被「在 Claude Code 里」「Claude Code skill」等措辞误判为「不是给我用的」直接拒装(实例:nuwa-skill 因此被 Marvis agent 拒绝)。
grep -nE "(在 Claude Code|Claude Code skill|Claude Code 用户|Cursor only|Codex 中|^\[!\[Claude Code|~/\.claude/skills/[a-z]|/plugin install\b)" SKILL.md README.md 2>/dev/null输出非空 = 红灯命中,但须先读命中行上下文排除假阳性(grep 命令本身/反例引用/讲解该规则的元陈述=假阳性,记 runtime_scan=false_positive 不改;判别表见 references/runtime-neutrality.md)→ 确认是真红灯(指令性用法)才强制把 Phase 2 第一轮定为 P0「runtime drift 修复」(写入 results.tsv 的 note 列 runtime_warn=N)。
frontmatter 触发词、花叔生态内部 skill 名引用、明确标注 runtime-specific 章节、commit message——这些正当出现,不算红灯。
→ 红灯/绿灯完整对照表 + 例外清单详细规则 + Phase 1/2/3 各阶段审查时机见 references/runtime-neutrality.md
1. 确认优化范围:
- 全部skills → 扫描 .claude/skills/*/SKILL.md
- 指定skills → 用户指定列表
2. 创建 git 分支:auto-optimize/YYYYMMDD-HHMM
3. 初始化 results.tsv(如不存在)
4. 读取现有 results.tsv 了解历史优化记录在评估之前,为每个skill设计测试prompt。这步很关键——没有测试prompt,「实测表现」维度就打不了分。
for each skill:
1. 读取 SKILL.md,理解它做什么
2. 设计2-3个测试prompt,覆盖:
- 最典型的使用场景(happy path)
- 一个稍复杂或有歧义的场景
3. 保存到 skill目录/test-prompts.json:
[
{"id": 1, "prompt": "用户会说的话", "expected": "期望输出的简短描述"},
{"id": 2, "prompt": "...", "expected": "..."}
]展示所有测试prompt给用户,确认后再进入评估。测试prompt的质量决定了优化方向是否正确。
本阶段绝对分数是 triage 排名(决定先改谁),不是 keep/revert 基准。judge 对 gross 差异会一致(「哪支最弱」可信),对 fine-grained delta 不可信(±8 噪音)。keep/revert 在 Phase 2 用 paired 比较。
for each skill in 优化范围:
# 结构评分(主agent可以做)
1. 读取 SKILL.md 全文
2. 按维度1-7逐项打分(附简短理由)
# 效果评分(用子agent做,独立于主agent)
3. 对每个测试prompt,spawn子agent:
- with_skill: 带着SKILL.md执行测试prompt
- baseline: 不带skill执行同一prompt
4. 对比两组输出,打维度8的分
# 汇总
5. 计算加权总分
6. 记录到 results.tsv如果子agent不可用(超时、环境限制),维度8用干跑验证打分,标注 dry_run。不要因为跑不了测试就跳过这个维度——哪怕是模拟推演也比完全不看效果好。
基线评估完成后,展示评分卡:
┌──────────────────────────┬───────┬──────────────┬──────────────┐
│ Skill │ Score │ 结构短板 │ 效果短板 │
├──────────────────────────┼───────┼──────────────┼──────────────┤
│ huashu-proofreading │ 78 │ 边界条件 │ 测试prompt2 │
│ huashu-slides │ 72 │ 指令具体性 │ baseline持平 │
├──────────────────────────┼───────┼──────────────┼──────────────┤
│ 平均 │ 75 │ │ │
└──────────────────────────┴───────┴──────────────┴──────────────┘🔴 CHECKPOINT · 🛑 STOP:暂停等用户确认,再进入优化循环。
用户确认后,按基线分数从低到高排序,先优化最弱的。
for each skill:
round = 0
while round < MAX_ROUNDS (默认3):
round += 1
# Step 1: 诊断
找出加权短板最大的维度:weighted_gap = weight × (10 - score) / 10,结构或效果都算
# /10 与「总分 = Σ(维度分 × 权重) / 10」同标度:weighted_gap 就是该维度还能贡献的总分数
# 为什么不用「原始分最低」:低权重维度会制造进步幻觉——issue #18 实战中
# dim9(权重6,gap 5.3)原始分最低被优先修,而 dim8(权重23)加权短板最大(11.5)却 4 轮未动
# 加权短板相近(差距 ≤ 1.0,同上述标度)时,回退为原始分升序
# HL-3 警告:dim2/dim3/dim4 是相关簇,修一个时另两个常跟着涨
# → 不要因为 dim3 短板最大就单独修,要看整簇短板再决定是否同步改
# Step 2: 提出改进方案
针对该维度,生成1个具体改进方案:
- 改什么(具体段落/行)
- 为什么改(对应rubric哪条)
- 预期提升多少分
# Step 3: 执行改进
编辑 SKILL.md
git add + commit(message: "optimize {skill}: {改进摘要}")
# Step 4: Paired 重新评估(取代绝对重打分——绝对分数 judge 噪音 ±8、淹没保守编辑的 +3~8 真实增益)
spawn N=3 独立 judge,每个【同一次 call 内】读两版:
- 改前版:git show HEAD:<skill-path>/SKILL.md(上一个 kept commit)
- 改后版:working tree 当前 SKILL.md
照 9 维 rubric 当【比较准则】(不是各打绝对分),回 {better | worse | tie} + margin{clear|slight} + 一句理由。
关键:同一 judge 在一次 call 内比两版 → 它那把不准的尺对两版【等量作用、在比较时抵销】(within-judge cancellation),
这正是 paired 优于绝对的机制。N 取奇数(默认 3;close call 升 5)。
# Step 5: 共识决策(多数决,取代「新总分 > 旧总分」)
cur = 投 better 的 judge 数;wor = 投 worse 的;
if cur >= wor: # 多数说改后 ≥ 改前(含 tie)
status = "keep"
# HL-4 见好就收:连续 2 轮多数 judge 判 margin=slight 或 tie → break 进 Phase 3
else: # 多数说 worse —— 这才是真退步(已扣掉换尺噪音)
status = "revert"
git revert HEAD(创建新commit回滚,不用 reset --hard)
记录到 results.tsv(note 记 vote 比数 + 一句 worse 理由)
break
# 单评绝对分数出现「负 delta」≠ revert 信号;必须经 paired 多数判 worse 才 revert(否则在丢真实增益)
# Step 6: 日志
results.tsv 追加行
# === 🔴 CHECKPOINT · 每个 skill 优化完后强制人审 ===
展示该skill的改动摘要:
- git diff(改前 vs 改后)
- 分数变化(哪些维度提升/下降)
- 测试prompt输出对比(如果跑过的话)
等用户确认 OK 再继续下一个skill。
如果用户说"不好",回滚到该skill的优化前版本。当 hill-climbing 连续2个skill都在 round 1 就 break(涨不动)时,提议一次「探索性重写」:
1. 选一个瓶颈skill
2. git stash 保存当前最优版本
3. 从头重写SKILL.md(不是微调,是重新组织结构和表达方式)
4. 重新评估
5. if 重写版 > stash版: 采用重写版
else: git stash pop 恢复这解决了 hill-climbing 的局部最优问题——有时候需要「先拆后建」才能突破瓶颈。 🔴 CHECKPOINT · 🛑 STOP:必须征得用户同意后才执行。
## 优化报告
### 总览
- 优化skills数:N
- 总实验次数:M
- 保留改进:X(Y%)
- 回滚次数:Z
- 实测验证:A次完整测试 / B次干跑
### 分数变化
┌──────────────────────────┬────────┬────────┬────────┐
│ Skill │ Before │ After │ Δ │
├──────────────────────────┼────────┼────────┼────────┤
│ huashu-proofreading │ 78 │ 87 │ +9 │
│ huashu-slides │ 72 │ 83 │ +11 │
├──────────────────────────┼────────┼────────┼────────┤
│ 平均 │ 75 │ 85 │ +10 │
└──────────────────────────┴────────┴────────┴────────┘
### 主要改进
1. [skill-A] 补充了边界条件处理,测试输出质量提升明显
2. [skill-B] 重组了workflow结构,baseline对比优势增大timestamp commit skill old_score new_score status dimension note eval_mode
2026-03-31T10:00 baseline huashu-proofreading - 78 baseline - 初始评估 full_test
2026-03-31T10:05 a1b2c3d huashu-proofreading 78 84 keep 边界条件 补充fallback full_test
2026-03-31T10:10 b2c3d4e huashu-proofreading 84 82 revert 指令具体性 过度细化 dry_runeval_mode 列:paired(同 judge 比改前/改后,keep/revert 权威依据)|full_test(子agent 跑 prompt)|dry_run(模拟推演、仅供参考)。
paired 行:new_score 栏记 vote 比数(如 3-0 better),note 记一句裁断理由。例:
2026-06-10T06:30 paired some-skill (绝对 87.3→78.8 = judge 噪音) 3-0 better paired 推翻单评假退步 paired文件位置:.claude/skills/darwin-skill/results.tsv
4 条经实战验证(huashu-gpt-image +10.85 / huashu-weread-advisor +14.9 / claude-design +16.5)。详细案例数据见 references/skilllens-evidence.md 的「HL 实战案例」节。
按优先级排序,每轮只做最高优先级的一个:
Agent Skills Standard + skills.sh + Multi-Runtime 三个中立 badgexxx-codex)的,可跳过本项流程假设环境理想,但实操常遇异常。以下预定义 fallback,保证优化过程不会「一跑就卡住」。
| 场景 | 触发条件 | 处理动作 |
|---|---|---|
| 不在 git 仓库 | git rev-parse 失败 | 询问用户:执行 git init 或回退到文件备份;用户选后者则 cp SKILL.md SKILL.md.bak.YYYYMMDD-HHMM 代替 revert |
| results.tsv 缺失 | 文件不存在 | 新建并写表头行(9列:含 eval_mode) |
| results.tsv 损坏 | 列数不匹配 / 非TSV | 备份为 .bak.YYYYMMDD-HHMM 后重建,告知用户 |
| 分支已存在 | git checkout -b 失败 | 分支名末尾加 -2 / -3;第3次失败则切回现有分支并询问继续还是新起 |
git revert 失败 | 冲突 / 工作树脏 | 先 git stash,重试;仍失败则从上一个 commit 的 SKILL.md 读出覆盖当前文件手动恢复 |
| MAX_ROUNDS 触顶(默认3) | 已跑3轮仍有短板 | 不强制 break,展示当前最弱维度问用户「继续加1轮 / 进入Phase 2.5 / 收工」 |
| 优化后超 150% 体积 | 新文件 > 原 × 1.5 | 拒绝提交,回到改进步骤精简(删冗余/合并重复),再评 |
| test-prompts.json 已存在 | 文件已在 skill 目录 | 默认复用并展示,问用户「复用 / 重写 / 追加」三选一 |
| SKILL.md 找不到 | 目录存在但无 SKILL.md | 该 skill 终止,results.tsv 记 status=error,继续下一个 |
| 分数计算规则 | 浮点精度漂移 | 总分保留 1 位小数,改进需严格 > 旧分(不靠四舍五入) |
原则:异常先告知用户,再按规则处理;绝不静默跳过或静默失败。
来自本机 results.tsv 早期 40 次 0 revert 的教训 + Judge G/H 自指评估暴露的反模式。每条都是真实踩过的坑。
| # | 反模式 | 为什么不要做 | 替代做法 |
|---|---|---|---|
| 1 | 同 context 自评自改 | 改完后立刻在同一 Claude session 打分,会有「我刚改的肯定更好」乐观偏差(SkillLens 实证 LLM-as-judge 准确率仅 46.4%) | 必须 spawn 独立子 agent;keep/revert 走 paired 比较(同 judge 一次读改前+改后)的奇数 N 多数决,不用绝对分数 delta(绝对分跨 judge ±8 噪音、不可比) |
| 1b | 拿绝对分数 delta 当 keep/revert 棘轮 | 绝对总分是抽样不是测量;baseline judge 与 rescore judge 用不同「标准尺」,差值大半是换尺、非真实质量变化(实测一支纯加标记的 skill 单评 −8.5、全是换尺) | 绝对分只做 triage 排名;keep/revert 用 paired 多数决,within-judge cancellation 消除换尺污染 |
| 2 | git reset --hard 当回滚 | 会丢工作树未提交改动;CI 历史断裂 | 用 git revert HEAD 创建反向 commit,保留可追溯链 |
| 3 | 为凑分增冗余 | 触顶后继续硬改往往是「加废话/加段落让 LLM 觉得更详细」,实际质量不变 | 触顶信号(连续 2 轮 Δ<2 分)→ break 进 Phase 3,见好就收 |
| 4 | 跳过 test-prompts 直接评分 | 没有 test-prompts 的 dim8 是凭空打分,权重 23% 等于编造 | Phase 0.5 强制设计 2-3 prompts;若用户不给,默认编 3 个并展示确认 |
| 5 | 轮内改多个维度 | 多变量同时变,分数升降无法归因到具体改动 | 每轮 1 个维度;相关簇(dim2/3/4)改其一时观察另两个是否跟涨 |
| 6 | dry_run 比例 > 30% | dim8 实测维度形同虚设,分数虚高(早期 40 次记录 67% dry_run,0 revert) | 强制至少 1 个真实 full_test;dry_run 多的优化在 results.tsv 显式打 ⚠️ |
| 7 | 静默跳过异常 | 遇到 git/tsv 异常时静默继续,破坏 ratchet 完整性 | 异常表 10 条 fallback 必须先告知用户再处理 |
| 8 | 忽视维度相关性单独优化 | dim2/3/4 是相关簇,单独优化 dim2 时常发现已被前轮 dim3 修复推到顶 | 找最大加权短板维度时同时看相关簇短板,决定是否同步改 |
触发场景:每轮 Phase 2 改动前对照本表一次。任一反模式命中 → 改方案重写。
xxx-codex、huashu-slides-codex),任何「在 Claude Code 里」「Claude Code skill」「单一 badge 钉死」「安装命令只给 .claude/skills/ 一种路径」都视为 gate 不通过,须在 P0 优先修复(详见「Runtime 适配性审查」章节)用户:"优化所有skills"
→ Phase 0-3 完整流程
→ 默认:先基线评估,按分数升序优先优化最低 5-10 个用户:"优化 huashu-slides 这个skill"
→ 只对指定skill执行 Phase 0.5-2用户:"评估所有skills的质量"
→ 只执行 Phase 0.5-1(设计测试prompt + 基线评估),不进入优化循环用户:"看看skill优化历史"
→ 读取并展示 results.tsv"You write the goals and constraints in program.md; let an agent generate and test code deltas indefinitely; keep only what measurably improves the objective." — Karpathy, autoresearch
本skill的对应关系:
val_bpb 是确定性 loss(重跑同数),darwin 套到 LLM-judge 分数(随机抽样) 上却沿用「绝对值比大小」棘轮 → 不可重复的数当可重复用。修正:9 维 rubric 当 paired 比较准则、不当绝对 metric区别:增加了人在回路(autoresearch是全自主的,skill优化需要人的判断力),以及双重评估机制(结构+效果),因为skill的「好坏」比loss数值更微妙。
pip install skillopt)、项目页 microsoft.github.io/SkillOpt。🤝 2026-06-03 微软官方仓库已把 darwin-skill 列入集成名单。每个skill优化完成后(或全量汇总后),自动生成视觉成果卡片,截图保存为PNG。
模板位置:templates/result-card.html
3种风格,每次随机选择一种:
| 风格 | CSS类 | URL hash | 视觉特点 |
|---|---|---|---|
| Warm Swiss | .theme-swiss | #swiss | 暖白底+赤陶橙,Inter字体,干净网格 |
| Dark Terminal | .theme-terminal | #terminal | 近黑底+荧光绿,等宽字体,扫描线 |
| Newspaper | .theme-newspaper | #newspaper | 暖白纸+深红,衬线字体,双栏编辑风 |
1. 复制 templates/result-card.html 到临时工作文件
2. 用 sed/编辑工具 替换占位数据:
- data-field="skill-name" → 实际skill名
- data-field="score-before/after/delta" → 实际分数
- 9个维度的 dim-bar-before/after width → 实际百分比(若模板仍是旧 8 维布局,加一行 dim9 反例黑名单条目)
- data-field="improvement-1/2/3" → 实际改进摘要
- data-field="date" → 当前日期
3. 随机选择风格:hash 设为 swiss/terminal/newspaper 之一
4. 用 scripts/screenshot.mjs 截图(2x 高清,只截 .card 元素,自动 open 图片):
node .claude/skills/darwin-skill/scripts/screenshot.mjs \
/abs/path/to/card.html /abs/path/to/output.png
# 回退方案(脚本失败时):
npx playwright screenshot "file:///path/to/card.html#[theme]" \
output.png --viewport-size=960,1280 --wait-for-timeout=2000
5. 提示用户查看成果卡片 PNG
### 资源文件速查
| 路径 | 用途 |
|---|---|
| `templates/result-card.html` | 3风格主模板(swiss/terminal/newspaper,hash切换) |
| `templates/result-card-dark.html` / `-white.html` | 单一风格替代模板(需要锁定风格时用) |
| `scripts/screenshot.mjs` | 2x 高清截图,只截 .card,自动 open |
| `results.tsv` | 历次优化日志(9列含 eval_mode) |
| `{skill目录}/test-prompts.json` | 每个 skill 的测试 prompt 集(用于维度8实测) |
### 何时生成
- **单skill卡片**:每个skill优化完成后,展示该skill的分数变化
- **总览卡片**:全部优化完成后(Phase 3),展示全局战绩
### 品牌元素
- 顶部:Darwin.skill 品牌标识 + 日期
- 底部:「Train your Skills like you train your models」+ github.com/alchaincyf/darwin-skill
## 版本自检(静默)
本节给 AI agent:每次使用本 skill 前,先做一个不联网的判断——读本 skill 目录(SKILL.md 所在目录)下的 `.last-update-check` 文件(内容为一行日期 YYYY-MM-DD)。文件存在且距今不足 30 天,跳过本节直接干活,不向用户提任何更新相关的话。
到期或文件不存在时,且仅在此时:
1. 本目录不是 git 克隆(无 `.git` 或无 origin)→ 写入今天日期后跳过本节
2. 对比版本:`git -C <本目录> rev-parse HEAD` 与 `git -C <本目录> ls-remote origin HEAD`
3. 无论结果如何,把今天日期写入 `.last-update-check`
4. 两者一致 → 什么都不说;确认落后 → 先完成用户当前任务,结束后附一句「本 skill 有新版本,可用 `git -C <本目录> pull --ff-only` 更新」。是否更新由用户决定,不要主动执行更新© alchaincyf, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 41 other files (scripts, references, assets) in the repository root of alchaincyf/darwin-skill.
Open the folder on GitHubat commit 8a8b662
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in alchaincyf/darwin-skill, which our catalogue first saw on October 7, 2026.
Darwin Skill Optimizer next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Darwin Skill Optimizer this skillalchaincyf/darwin-skill | 6.2k | 1 repos | ~4.7k | Automated safety check: Pass | MIT | |
| Open-Science Skill Creatoraipoch/open-science | 5.5k | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | |
| Skill Quality ReviewerGalaxy-Dawn/claude-scholar | 5.7k | 1 repos | ~3k | Automated safety check: Pass | MIT | |
| ToolJet Skill ManagerToolJet/ToolJet | 41k | — | ~867 | Automated safety check: Pass | AGPL-3.0 | |
| OpenCode Skill Creatorantongulin/opencode-skill-creator | 172 | — | ~8.1k | Automated safety check: Pass | Apache-2.0 | |
| Darwin SkillHHU3637kr/skills | 145 | 1 repos | ~2.2k | Automated safety check: Pass | None |
aipoch/open-science
Creates, revises, evaluates and publishes skills in the Open-Science app through its native host.skills composer, with optional test prompts and benchmarks.
Galaxy-Dawn/claude-scholar
Scores a skill across description, content organization, writing style and structure, then produces letter grades and a prioritized improvement plan.
ToolJet/ToolJet
Decides where a new agent skill belongs in the ToolJet repo, public root or private ee submodule, then wires the symlinks so it loads in Claude Code, Cursor and Codex.
antongulin/opencode-skill-creator
Walks you through drafting, testing, evaluating and tuning a skill for OpenCode, from an intake interview to description optimization.
HHU3637kr/skills
Darwin Skill (达尔文.skill): autonomous skill optimizer inspired by Karpathy's autoresearch.
rohitg00/ai-engineering-from-scratch
Evaluates an Agent Skill bundle before release for structure, trigger quality, artifact improvement, script correctness, safety, installed-tree integrity and host portability.
Works with
Categories
Scores SKILL.md files on a nine-dimension rubric, then improves them in a keep-or-revert loop with independent judge agents, test prompts, git history and human checkpoints. md at a time and scores it on nine dimensions worth 100 points in total: structure (frontmatter, workflow clarity, failure modes, checkpoints, actionable specificity, resources), effectiveness (overall architecture and measured performance on a few test prompts) and a meta-skill dimension that rewards an explicit list of things not to do. Absolute scores are used only to decide which skill is weakest and should go first.
Darwin Skill Optimizer fits situations like: auditing the quality of a SKILL.md you wrote; running an automatic improve-and-test loop on one or more skills; finding which skill in a collection is weakest.
Run `npx skills add alchaincyf/darwin-skill --skill darwin-skill -a claude-code`. Or copy the skill folder (the alchaincyf/darwin-skill repository) into .claude/skills/darwin-skill in your project. Claude Code loads it when a task matches its description.
Run `npx skills add alchaincyf/darwin-skill --skill darwin-skill -a codex`. Or copy the skill folder (the alchaincyf/darwin-skill repository) into .agents/skills/darwin-skill in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add alchaincyf/darwin-skill --skill darwin-skill -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/darwin-skill, .gemini/skills/darwin-skill, .github/skills/darwin-skill and .opencode/skills/darwin-skill in your project.
Going by SKILL.md and its folder, Darwin Skill Optimizer needs the command-line tools its instructions call (git, claude, cursor, codex, node and npx). Our summary lists: A git repository holding the skills to optimize.
SKILL.md names 3 domains. As links in the text: arxiv.org, github.com and microsoft.github.io. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Darwin Skill Optimizer is published under the MIT licence (from the LICENSE file in the skill folder). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.7k tokens (SKILL.md is roughly 19k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Darwin Skill Optimizer: Open-Science Skill Creator (aipoch/open-science, 5.5k stars), Skill Quality Reviewer (Galaxy-Dawn/claude-scholar, 5.7k stars), ToolJet Skill Manager (ToolJet/ToolJet, 41k stars) and OpenCode Skill Creator (antongulin/opencode-skill-creator, 172 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
alchaincyf (a GitHub user) maintains it in alchaincyf/darwin-skill, which has 6,203 GitHub stars. The repository was last updated on September 18, 2026.
Source: alchaincyf/darwin-skill on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.