Darwin Skill Optimizer
alchaincyf/darwin-skill
Scores SKILL.md files on a nine-dimension rubric, then improves them in a keep-or-revert loop with independent judge agents, test prompts, git history and human checkpoints.
Runs and interprets a skill's Caliper eval: how often it succeeds over repeated attempts, whether it triggers at all, and whether it beats the agent without it.
$ npx skills add edonadei/caliper --skill evaluate-skill -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install edonadei/caliper evaluate-skill --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/edonadei/caliper.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/evaluate-skill .claude/skills/evaluate-skill && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "evaluate-skill" agent skill from https://github.com/edonadei/caliper/tree/main/skills/evaluate-skill into .claude/skills/evaluate-skill/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evaluate-skill", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/edonadei/caliper/tree/main/skills/evaluate-skillType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add edonadei/caliper --skill evaluate-skill -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install edonadei/caliper evaluate-skill --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/edonadei/caliper.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/evaluate-skill .agents/skills/evaluate-skill && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "evaluate-skill" agent skill from https://github.com/edonadei/caliper/tree/main/skills/evaluate-skill into .agents/skills/evaluate-skill/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evaluate-skill", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add edonadei/caliper --skill evaluate-skill -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install edonadei/caliper evaluate-skill --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/edonadei/caliper.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/evaluate-skill .cursor/skills/evaluate-skill && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "evaluate-skill" agent skill from https://github.com/edonadei/caliper/tree/main/skills/evaluate-skill into .cursor/skills/evaluate-skill/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evaluate-skill", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/edonadei/caliper.git --path skills/evaluate-skill--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add edonadei/caliper --skill evaluate-skill -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install edonadei/caliper evaluate-skill --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/edonadei/caliper.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/evaluate-skill .gemini/skills/evaluate-skill && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "evaluate-skill" agent skill from https://github.com/edonadei/caliper/tree/main/skills/evaluate-skill into .gemini/skills/evaluate-skill/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evaluate-skill", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install edonadei/caliper evaluate-skillInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add edonadei/caliper --skill evaluate-skill -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/edonadei/caliper.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/evaluate-skill .github/skills/evaluate-skill && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "evaluate-skill" agent skill from https://github.com/edonadei/caliper/tree/main/skills/evaluate-skill into .github/skills/evaluate-skill/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evaluate-skill", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add edonadei/caliper --skill evaluate-skill -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install edonadei/caliper evaluate-skill --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/edonadei/caliper.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/evaluate-skill .opencode/skills/evaluate-skill && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "evaluate-skill" agent skill from https://github.com/edonadei/caliper/tree/main/skills/evaluate-skill into .opencode/skills/evaluate-skill/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evaluate-skill", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
evaluate-skillRuns and interprets a skill's Caliper eval: how often it succeeds over repeated attempts, whether it triggers at all, and whether it beats the agent without it.
The skill operates the caliper command-line tool against an .eval.yaml spec and traces each failure to the place that fixes it. It separates four questions that a single score would blur: whether the skill fires when it should, whether it works once it fires, whether it earns its place against a control run without it, and whether it holds up across edits over time. Each has its own measurement and its own fix location, such as the skill's description, its body, the tasks, or the edit that moved the result.
When writing a spec, every task that checks results also checks activation, and prompts read like a real user's request with the skill left unnamed. The engine is chosen at run time and can differ for the skill and the judge, with claude-code, codex, pi and hermes available, and each attempt runs in a fresh empty working directory. Validation never touches the network, and a first run with one attempt is meant to shake out spec and harness errors before any score is trusted. A REFERENCE.md file covers the full spec format.
4 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit f3f0355. It shows what the files ask for, not the result of running them.
Pre-approves these tools, so the agent can use them without asking each time:
BashFrom allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pipxFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use pipx, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Evaluate Skill with Caliper loads about 1.9k tokens when it runs. Until then it costs about 76 tokens; SKILL.md has 1,029 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
allowed-tools: BashAutomated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from edonadei/caliper at commit f3f0355, republished under its MIT licence (© edonadei). 1,029 words, ~1,897 tokens.
.claude/skills/evaluate-skill/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.Operate Caliper: run a skill's eval, read what the results say, and trace each failure to the place that fixes it.
The caliper CLI must be on PATH. This skill can be copied into an agent without the Caliper repo, so install it if missing:
pipx install caliper-evalcaliper <command> --help is the authority on flags. REFERENCE.md covers the full spec format, engines, and what the results mean.
A skill eval answers four separate questions. Each has its own measurement and its own fix location, so keep them apart: a blended number hides which half broke.
| Question | Measured by | A failure is fixed in |
|---|---|---|
| Fires: does the agent reach for the skill when it should, and only then? | activates: on tasks, and trigger probes | the skill's description frontmatter |
| Works: once it fires, does it get the job done? | expect: / assert:, scored as the success rate | the skill's body |
| Earns: does it beat the agent without it? | the control (--ablate <skill-name>: the declared neighbourhood minus this skill) and caliper compare. For a truly bare agent, ablate every declared skill and every declared mcp: server, and isolate the run | the tasks: if the control passes too, the task is too easy |
| Holds: does it stay good across edits and over time? | caliper compare of each full run against the previous one, skill drift | the edit that moved it |
skills:
- ./SKILL.md # relative to the spec; installed, never preloaded
tasks:
- name: What success looks like
prompt: <what a real user would type; never names the skill>
expect: <natural-language pass/fail criterion>
assert: | # optional deterministic Python check
assert ...
activates: [my-skill] # frontmatter name: of the skill that should fire
- name: Unrelated work stays unrelated
prompt: <work no declared skill should answer>
activates: [] # a trigger probe: no judge, cheapWhen you write a spec, every task with expect: or assert: also asserts activates:, refusals included: the skill, plus any declared skill it delegates to on that task, and every prompt reads like a real user's request with the skill left unnamed.
The spec has no backend/model or judge: block. The engine is chosen at run time, independently for the skill and the judge: caliper run <spec> --model codex runs and grades on codex, and --judge-model picks a different judge. Backends are claude-code (default), codex, pi, and hermes. Each attempt runs in a fresh, empty workdir, so setup: builds fixtures there with relative paths.
caliper validate <spec>. It never touches the network.caliper run <spec> --k 1 to shake out spec and harness errors. Fix those before reading any score.caliper run <spec> --k 3 --ablate <skill-name>, once (write skill:<skill-name> if an mcp: server shares the name). This is the control: the declared neighbourhood without this skill. The skill isn't installed, so editing SKILL.md can't move its number. Keep its results path (caliper list <spec-name> marks ablated runs) and re-diff against it. Re-run it only when the tasks or the declared skills change.caliper run <spec> --k 3, then caliper compare <control.json> <spec-name> to see whether it earns its place. After each edit, also compare against the previous full run's path to see whether the edit held: a skill that got worse can still beat the control. A bare spec name resolves to that spec's latest run.The success rate (successes / usable) is the headline. pass@k and pass^k are secondary views under --verbose. Activation is a separate scoreboard, over a different population: report it beside the success rate, never averaged into it.
Trace every failing task to where its fix belongs:
| Signal | Where the fix belongs |
|---|---|
Run exits 2 (backend misconfigured, unavailable model, failed hook or MCP server, or every attempt unusable) | The environment or the task's hooks. Report it as a configuration problem, not a task failure |
⊘ unusable attempts (infra_error, timeout, judge_error) | Not the skill: rate limits, auth, or the judge. They are excluded from the score; re-run |
cheat outcome | The task: it leaks its answer. Tighten the task or sandbox: |
| Activation fails: an expected skill didn't fire | That skill's description |
| Activation fails: an unexpected skill fired too | The extra skill's description, or the overlap between the two |
Activation fails because one of the user's own skills fired (skill: in the report header) | Not the description alone: that skill is real competition in their setup. Decide with the user whether to sharpen the description or isolate the run |
| Activation passes, score low | The skill's body |
The judge's reasoning shows expect: was ambiguous | The task's grading: make the criterion observable, or add an assert: |
| Full run ≈ control, and the skill fired in the full run | The task: it doesn't need the skill. If the skill never fired, the fix is its description (above) |
At k=3 one attempt is a 33-point swing. Before calling a change a win or a regression, re-run at k≥5. caliper compare flags any drop and never gates. A rewritten or shortened skill is safe to ship when, at k≥5, it stays within about 5% of the previous score and still beats the control.
Runs load the user's own customizations by default (user skills, plugins, rules, settings and connectors; see REFERENCE.md for backend exceptions), which answers "does my skill work in my agent?". Isolate (--no-user-customizations, or user_customizations: false in the spec) when comparing backends or models, when the number leaves this machine (shared, published, compared with someone else's run), or when measuring the bare agent: each setup is different, so otherwise part of the delta is the setups. --ablate of the user's own skill needs no isolation, since both runs load the same setup.
Always tell the user which mode ran and what it loaded, from the report header's user customizations: line (absent means isolated), and relay any fix caliper compare suggests about it.
If the user hasn't decided what to test (a skill with no .eval.yaml, or an eval whose gaps they want found), suggest grill-skill: it interviews them and writes the spec. To design tasks yourself, follow "Designing evals" in REFERENCE.md.
assert:;activates:, and the spec has at least one trigger probe;caliper validate;.eval.yaml beside SKILL.md. Saved runs under .caliper/results/ are useful for diffing over time and safe to gitignore.© edonadei, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files in skills/evaluate-skill of edonadei/caliper.
Open the folder on GitHubat commit f3f0355
Evaluate Skill with Caliper next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Evaluate Skill with Caliper this skilledonadei/caliper | 206 | — | ~1.9k | Automated safety check: Notes | MIT | |
| Darwin Skill Optimizeralchaincyf/darwin-skill | 6.2k | 1 repos | ~4.7k | Automated safety check: Pass | MIT | |
| Skill Release Gaterohitg00/ai-engineering-from-scratch | 65k | — | ~1k | Automated safety check: Pass | MIT | |
| Open-Science Skill Creatoraipoch/open-science | 5.4k | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | |
| Skill JudgeshareAI-lab/Kode-CLI | 5.2k | 4 repos | ~7.5k | Automated safety check: Pass | Apache-2.0 | |
| Skill Quality ReviewerGalaxy-Dawn/claude-scholar | 5.7k | 1 repos | ~3k | Automated safety check: Pass | MIT |
alchaincyf/darwin-skill
Scores SKILL.md files on a nine-dimension rubric, then improves them in a keep-or-revert loop with independent judge agents, test prompts, git history and human checkpoints.
rohitg00/ai-engineering-from-scratch
Evaluates an Agent Skill bundle before release for structure, trigger quality, artifact improvement, script correctness, safety, installed-tree integrity and host portability.
aipoch/open-science
Creates, revises, evaluates and publishes skills in the Open-Science app through its native host.skills composer, with optional test prompts and benchmarks.
shareAI-lab/Kode-CLI
Evaluates the design quality of an agent skill against official specifications and patterns from existing examples, scoring it and suggesting improvements.
Galaxy-Dawn/claude-scholar
Scores a skill across description, content organization, writing style and structure, then produces letter grades and a prioritized improvement plan.
antongulin/opencode-skill-creator
Walks you through drafting, testing, evaluating and tuning a skill for OpenCode, from an intake interview to description optimization.
edonadei/caliper
Runs caliper's smoke evals against the real agent CLIs after a harness or MCP change, with a dry-run plan, failure triage and a report to attach to the PR.
edonadei/caliper
Interviews you about what a skill should do, writes an eval spec from your answers, then loops through run, diagnose and improve until the skill is ready to ship.
Categories
Runs and interprets a skill's Caliper eval: how often it succeeds over repeated attempts, whether it triggers at all, and whether it beats the agent without it. yaml spec and traces each failure to the place that fixes it. It separates four questions that a single score would blur: whether the skill fires when it should, whether it works once it fires, whether it earns its place against a control run without it, and whether it holds up across edits over time.
Evaluate Skill with Caliper fits situations like: running an eval for a skill you wrote and reading the results; finding out whether a skill triggers on the prompts it should; checking whether a skill beats the agent without it; comparing eval runs after editing a skill.
Run `npx skills add edonadei/caliper --skill evaluate-skill -a claude-code`. Or copy the skill folder (skills/evaluate-skill in edonadei/caliper) into .claude/skills/evaluate-skill in your project. Claude Code loads it when a task matches its description.
Run `npx skills add edonadei/caliper --skill evaluate-skill -a codex`. Or copy the skill folder (skills/evaluate-skill in edonadei/caliper) into .agents/skills/evaluate-skill in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add edonadei/caliper --skill evaluate-skill -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/evaluate-skill, .gemini/skills/evaluate-skill, .github/skills/evaluate-skill and .opencode/skills/evaluate-skill in your project.
Going by SKILL.md and its folder, Evaluate Skill with Caliper needs the command-line tools its instructions call (pipx). Our summary lists: The caliper CLI, installed with pipx install caliper-eval; Bash access. Its frontmatter pre-approves these tools: Bash.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.
Evaluate Skill with Caliper is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.9k tokens (SKILL.md is roughly 7.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Evaluate Skill with Caliper: Darwin Skill Optimizer (alchaincyf/darwin-skill, 6.2k stars), Skill Release Gate (rohitg00/ai-engineering-from-scratch, 65k stars), Open-Science Skill Creator (aipoch/open-science, 5.4k stars) and Skill Judge (shareAI-lab/Kode-CLI, 5.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
edonadei (a GitHub user) maintains it in edonadei/caliper, which has 206 GitHub stars. The repository holds 3 skills in this directory. The repository was last updated on October 5, 2026.
Source: edonadei/caliper on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.