Darwin Skill Optimizer
alchaincyf/darwin-skill
Scores SKILL.md files on a nine-dimension rubric, then improves them in a keep-or-revert loop with independent judge agents, test prompts, git history and human checkpoints.
Interviews you about what a skill should do, writes an eval spec from your answers, then loops through run, diagnose and improve until the skill is ready to ship.
$ npx skills add edonadei/caliper --skill grill-skill -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install edonadei/caliper grill-skill --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/edonadei/caliper.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/grill-skill .claude/skills/grill-skill && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "grill-skill" agent skill from https://github.com/edonadei/caliper/tree/main/skills/grill-skill into .claude/skills/grill-skill/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "grill-skill", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/edonadei/caliper/tree/main/skills/grill-skillType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add edonadei/caliper --skill grill-skill -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install edonadei/caliper grill-skill --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/edonadei/caliper.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/grill-skill .agents/skills/grill-skill && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "grill-skill" agent skill from https://github.com/edonadei/caliper/tree/main/skills/grill-skill into .agents/skills/grill-skill/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "grill-skill", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add edonadei/caliper --skill grill-skill -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install edonadei/caliper grill-skill --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/edonadei/caliper.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/grill-skill .cursor/skills/grill-skill && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "grill-skill" agent skill from https://github.com/edonadei/caliper/tree/main/skills/grill-skill into .cursor/skills/grill-skill/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "grill-skill", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/edonadei/caliper.git --path skills/grill-skill--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add edonadei/caliper --skill grill-skill -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install edonadei/caliper grill-skill --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/edonadei/caliper.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/grill-skill .gemini/skills/grill-skill && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "grill-skill" agent skill from https://github.com/edonadei/caliper/tree/main/skills/grill-skill into .gemini/skills/grill-skill/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "grill-skill", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install edonadei/caliper grill-skillInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add edonadei/caliper --skill grill-skill -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/edonadei/caliper.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/grill-skill .github/skills/grill-skill && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "grill-skill" agent skill from https://github.com/edonadei/caliper/tree/main/skills/grill-skill into .github/skills/grill-skill/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "grill-skill", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add edonadei/caliper --skill grill-skill -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install edonadei/caliper grill-skill --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/edonadei/caliper.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/grill-skill .opencode/skills/grill-skill && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "grill-skill" agent skill from https://github.com/edonadei/caliper/tree/main/skills/grill-skill into .opencode/skills/grill-skill/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "grill-skill", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
grill-skillInterviews you about what a skill should do, writes an eval spec from your answers, then loops through run, diagnose and improve until the skill is ready to ship.
This skill uses the caliper CLI, installed with pipx as caliper-eval, to build evaluations for other skills. It frames an eval as four questions: does the agent pick the skill when it should and only then, does the skill do the job once it fires, does it beat the same agent without it, and does it stay good across edits. Each question has its own kind of check, and the fixes land in the skill's description, its body or the task.
You start it with /grill-skill and an optional path to a SKILL.md, or it looks in the current folder and confirms. It first summarizes the skill and waits for your confirmation, then looks for an eval YAML file beside the skill: none means a new eval, one found means filling gaps. The interview asks one question at a time about the happy path, an edge case, adversarial input and neighboring skills, plus an unrelated silence probe. The spec is written only after the interview. REFERENCE.md holds commands and task-writing rules, and the excerpt is cut off before the later phases.
5 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 586b0df. It shows what the files ask for, not the result of running them.
Pre-approves these tools, so the agent can use them without asking each time:
BashReadWriteEditFrom allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pipxFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use pipx, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Grill Skill Eval Designer loads about 2.1k tokens when it runs. Until then it costs about 64 tokens; SKILL.md has 1,253 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
allowed-tools: Bash, Read, Write, EditAutomated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from edonadei/caliper at commit 586b0df, republished under its MIT licence (© edonadei). 1,253 words, ~2,058 tokens.
.claude/skills/grill-skill/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.Interview the user to design a skill's eval, then loop run → diagnose → improve until it ships. Requires caliper (pipx install caliper-eval if missing). Commands, spec skeleton, and task-writing rules: REFERENCE.md.
An eval answers four questions, and the interview covers each:
activates: and trigger probes. Fixed in the description.expect: / assert:. Fixed in the body./grill-skill [path]: optional path to a SKILL.md.
SKILL.md in the cwd; if found, confirm before proceeding, else ask where it is.Read the SKILL.md. Summarize what it does, when it triggers, and what a successful run looks like. Ask the user to confirm your reading. Wait for confirmation before continuing.
Look for *.eval.yaml beside the SKILL.md (try <dir-name>.eval.yaml first).
Interview one question at a time and wait for each answer. Build every task from the user's answers, and write the spec only after the interview.
Elicit, one question at a time:
skills: and add a neighbour probe: a prompt that belongs to the neighbour, with activates: [<neighbour>].Always propose a silence probe as well: unrelated work, activates: [].
Turn each answer into a task: a realistic prompt that never names the skill, an observable expect, an assert when the outcome is checkable, and activates: on execution tasks, naming the skill plus any declared skill it delegates to on that task, so the run can tell a description failure from a body failure. The harness is single-shot: nobody answers the agent's questions. If the skill asks before acting, the task judges that first turn: expect: the question, assert: nothing was done yet. Show the proposed YAML and confirm before writing.
Write the spec beside SKILL.md, named <dir-name>.eval.yaml, with skills: [./SKILL.md] and no engine: the backend and model are picked at run time with --model / --judge-model, so if the SKILL.md targets a non-default agent, tell the user which flag to pass. Then tell the user to commit the spec now, beside SKILL.md.
Read the existing spec and report its tasks, grouped by the question each answers. Name any question with no task: execution tasks without activates:, no trigger probe, no deterministic assert:. Ask what behaviors are missing or under-tested before proposing or writing anything, even if the user only asked you to inspect it. Sharpen each gap into a task, show it, and confirm before writing it in.
Runs load the user's own customizations by default (user skills, plugins, rules, settings and connectors; see REFERENCE.md for backend exceptions), which answers "does my skill work in my agent?". Isolate (--no-user-customizations, or user_customizations: false in the spec) when comparing backends or models, when the number leaves this machine (shared, published, compared with someone else's run), or when measuring the bare agent: each setup is different, so otherwise part of the delta is the setups. --ablate of the user's own skill needs no isolation, since both runs load the same setup.
Always tell the user which mode ran and what it loaded, from the report header's user customizations: line (absent means isolated), and relay any fix caliper compare suggests about it.
Validate the spec, then run at k=1 (commands in REFERENCE.md). Show the results. Fix any harness or config error (not a task failure) before moving on. If a run stops with Not logged in, ask the user whether to run the login command the error names. With their consent, run it (it opens a browser for them to finish) and rerun; never run it without consent. hermes model is an interactive picker, so ask the user to run it in their own terminal instead.
Before anyone edits the skill, check that the tasks can tell it apart from the control: the declared neighbourhood with this skill removed. Run the control (--ablate <skill-name>, or skill:<skill-name> if an mcp: server shares the name) and a full run, both at k=3, then caliper compare them. Check activation first: if the skill never fired in the full run, both runs measure the same agent, and the fix is its description (Phase 5), not the tasks.
Keep the control's results path (caliper list <spec-name> shows which run was ablated). The skill isn't installed in that run, so editing SKILL.md can't move its number: re-diff against it instead of re-running it. Re-run it only when the tasks or the declared skills change.
Read each failing task before suggesting a fix, and say where the fix belongs:
| Signal | Where the fix belongs |
|---|---|
Run exits 2 (backend misconfigured, unavailable model, failed hook or MCP server, or every attempt unusable) | The environment or the task's hooks, not the skill. Fix and re-run |
⊘ unusable attempts (infra_error, timeout, judge_error) | Not the skill: rate limits, auth, or the judge. Fix and re-run |
cheat outcome | The task: it leaks its answer. Tighten the task or sandbox: |
| Activation fails: an expected skill didn't fire | That skill's description |
| Activation fails: an unexpected skill fired too | The extra skill's description, or the overlap between the two |
Activation fails because one of the user's own skills fired (skill: in the report header) | Not the description alone: that skill is real competition in their setup. Decide with the user whether to sharpen the description or isolate the run |
| Activation passes, score low | The skill's body |
The judge's reasoning shows expect: was ambiguous | The task's grading: make the criterion observable, or add an assert: |
| Full run ≈ control, and the skill fired in the full run | The task (see Phase 4). If it never fired, the description |
Then ask whether to iterate or finish.
SKILL.md, re-run at k=3 and caliper compare it twice: against the previous full run (did the edit hold?) and against the kept control (does it still earn its place?). A skill that got worse can still beat the control. Diagnose again. At k=3 one attempt is a 33-point swing, so confirm a surprising win or loss at k≥5 before acting on it. Loop.SKILL.md and the .eval.yaml together.© edonadei, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files in skills/grill-skill of edonadei/caliper.
Open the folder on GitHubat commit 586b0df
Grill Skill Eval Designer next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Grill Skill Eval Designer this skilledonadei/caliper | 208 | — | ~2.1k | Automated safety check: Notes | MIT | |
| Darwin Skill Optimizeralchaincyf/darwin-skill | 6.2k | 1 repos | ~4.7k | Automated safety check: Pass | MIT | |
| Skill Release Gaterohitg00/ai-engineering-from-scratch | 66k | — | ~1k | Automated safety check: Pass | MIT | |
| Open-Science Skill Creatoraipoch/open-science | 5.5k | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | |
| Skill JudgeshareAI-lab/Kode-CLI | 5.2k | 4 repos | ~7.5k | Automated safety check: Pass | Apache-2.0 | |
| Skill Quality ReviewerGalaxy-Dawn/claude-scholar | 5.7k | 1 repos | ~3k | Automated safety check: Pass | MIT |
alchaincyf/darwin-skill
Scores SKILL.md files on a nine-dimension rubric, then improves them in a keep-or-revert loop with independent judge agents, test prompts, git history and human checkpoints.
rohitg00/ai-engineering-from-scratch
Evaluates an Agent Skill bundle before release for structure, trigger quality, artifact improvement, script correctness, safety, installed-tree integrity and host portability.
aipoch/open-science
Creates, revises, evaluates and publishes skills in the Open-Science app through its native host.skills composer, with optional test prompts and benchmarks.
shareAI-lab/Kode-CLI
Evaluates the design quality of an agent skill against official specifications and patterns from existing examples, scoring it and suggesting improvements.
Galaxy-Dawn/claude-scholar
Scores a skill across description, content organization, writing style and structure, then produces letter grades and a prioritized improvement plan.
antongulin/opencode-skill-creator
Walks you through drafting, testing, evaluating and tuning a skill for OpenCode, from an intake interview to description optimization.
edonadei/caliper
Runs and interprets a skill's Caliper eval: how often it succeeds over repeated attempts, whether it triggers at all, and whether it beats the agent without it.
edonadei/caliper
Runs caliper's smoke evals against the real agent CLIs after a harness or MCP change, with a dry-run plan, failure triage and a report to attach to the PR.
Categories
Interviews you about what a skill should do, writes an eval spec from your answers, then loops through run, diagnose and improve until the skill is ready to ship. This skill uses the caliper CLI, installed with pipx as caliper-eval, to build evaluations for other skills. It frames an eval as four questions: does the agent pick the skill when it should and only then, does the skill do the job once it fires, does it beat the same agent without it, and does it stay good across edits.
Grill Skill Eval Designer fits situations like: designing the first eval for a skill that has none; finding the cases an existing skill eval does not cover; checking that a skill triggers on the right prompts and stays quiet otherwise; comparing a skill against the same agent without it.
Run `npx skills add edonadei/caliper --skill grill-skill -a claude-code`. Or copy the skill folder (skills/grill-skill in edonadei/caliper) into .claude/skills/grill-skill in your project. Claude Code loads it when a task matches its description.
Run `npx skills add edonadei/caliper --skill grill-skill -a codex`. Or copy the skill folder (skills/grill-skill in edonadei/caliper) into .agents/skills/grill-skill in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add edonadei/caliper --skill grill-skill -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/grill-skill, .gemini/skills/grill-skill, .github/skills/grill-skill and .opencode/skills/grill-skill in your project.
Going by SKILL.md and its folder, Grill Skill Eval Designer needs the command-line tools its instructions call (pipx). Our summary lists: The `caliper` CLI (`pipx install caliper-eval`); A SKILL.md file to build an eval for. Its frontmatter pre-approves these tools: Bash, Read, Write, Edit.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.
Grill Skill Eval Designer is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.1k tokens (SKILL.md is roughly 8.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Grill Skill Eval Designer: Darwin Skill Optimizer (alchaincyf/darwin-skill, 6.2k stars), Skill Release Gate (rohitg00/ai-engineering-from-scratch, 66k stars), Open-Science Skill Creator (aipoch/open-science, 5.5k stars) and Skill Judge (shareAI-lab/Kode-CLI, 5.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
edonadei (a GitHub user) maintains it in edonadei/caliper, which has 208 GitHub stars. The repository holds 3 skills in this directory. The repository was last updated on October 9, 2026.
Source: edonadei/caliper on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.