Agent skill

Grill Skill Eval Designer

by edonadei in edonadei/caliper

Interviews you about what a skill should do, writes an eval spec from your answers, then loops through run, diagnose and improve until the skill is ready to ship.

MITAuto-check: notesAgent Workflows

Install Grill Skill Eval Designer

skills CLI
$ npx skills add edonadei/caliper --skill grill-skill -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install edonadei/caliper grill-skill --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/edonadei/caliper.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/grill-skill .claude/skills/grill-skill && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
grill-skill
GitHub stars
208
Token cost
~2.1k tokens
SKILL.md length
1,253 words
Files
3
Skills in repo
3
Repo updated
First seen
Licence
MIT

At a glance

Interviews you about what a skill should do, writes an eval spec from your answers, then loops through run, diagnose and improve until the skill is ready to ship.

  • Works in 5 steps: Understand → Detect eval mode → First run → …
  • Designing the first eval for a skill that has none
  • SKILL.md covers Entry point, Phase 1 — Understand, Phase 2 — Detect eval mode and Whose setup is measured, plus 3 more sections
  • Calls pipx

What it does

This skill uses the caliper CLI, installed with pipx as caliper-eval, to build evaluations for other skills. It frames an eval as four questions: does the agent pick the skill when it should and only then, does the skill do the job once it fires, does it beat the same agent without it, and does it stay good across edits. Each question has its own kind of check, and the fixes land in the skill's description, its body or the task.

You start it with /grill-skill and an optional path to a SKILL.md, or it looks in the current folder and confirms. It first summarizes the skill and waits for your confirmation, then looks for an eval YAML file beside the skill: none means a new eval, one found means filling gaps. The interview asks one question at a time about the happy path, an edge case, adversarial input and neighboring skills, plus an unrelated silence probe. The spec is written only after the interview. REFERENCE.md holds commands and task-writing rules, and the excerpt is cut off before the later phases.

When your agent uses it

  • Designing the first eval for a skill that has none
  • Finding the cases an existing skill eval does not cover
  • Checking that a skill triggers on the right prompts and stays quiet otherwise
  • Comparing a skill against the same agent without it

Example prompts

  • “Run /grill-skill on skills/pdf-forms/SKILL.md and interview me about its eval.”
  • “My changelog skill has no eval; help me decide what it should test.”
  • “Look at the existing eval beside this SKILL.md and tell me what it misses.”
  • “Add a neighbour probe for the release-notes skill that this one might get confused with.”

Requirements

  • The `caliper` CLI (`pipx install caliper-eval`)
  • A SKILL.md file to build an eval for
  • Pre-approved tools (allowed-tools): Bash, Read, Write, Edit

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Understand
  2. Detect eval mode
  3. First run
  4. Check the tasks need the skill
  5. Diagnose and iterate

What it can do on your machine

Read from SKILL.md and the folder at commit 586b0df. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Bash
    • Read
    • Write
    • Edit

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pipx

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pipx, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Grill Skill Eval Designer loads about 2.1k tokens when it runs. Until then it costs about 64 tokens; SKILL.md has 1,253 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~64
When it runs · the whole SKILL.md, loaded when a task matches
~2.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Bash, Read, Write, Edit

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from edonadei/caliper at commit 586b0df, republished under its MIT licence (© edonadei). 1,253 words, ~2,058 tokens.

Download SKILL.mdSave it as .claude/skills/grill-skill/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
grill-skill
description
Interview the user to decide what a skill's eval should test, then build the spec and iterate until the skill ships. Use when the user wants help designing a skill's eval, has a skill with no eval, or wants to find what an existing eval misses.
allowed-tools
Bash, Read, Write, Edit

Grill Skill

Interview the user to design a skill's eval, then loop run → diagnose → improve until it ships. Requires caliper (pipx install caliper-eval if missing). Commands, spec skeleton, and task-writing rules: REFERENCE.md.

An eval answers four questions, and the interview covers each:

  • Fires: does the agent reach for the skill when it should, and only then? Tested with activates: and trigger probes. Fixed in the description.
  • Works: once it fires, does it do the job? Tested with expect: / assert:. Fixed in the body.
  • Earns: does it beat the agent without it? Tested against the control: the declared neighbourhood with this skill removed. A task the control passes too is fixed in the task.
  • Holds: does it stay good across edits? Tested by comparing each new full run against the previous one.

Entry point

/grill-skill [path]: optional path to a SKILL.md.

  • Path given: use it.
  • No path: look for SKILL.md in the cwd; if found, confirm before proceeding, else ask where it is.

Phase 1 — Understand

Read the SKILL.md. Summarize what it does, when it triggers, and what a successful run looks like. Ask the user to confirm your reading. Wait for confirmation before continuing.

Phase 2 — Detect eval mode

Look for *.eval.yaml beside the SKILL.md (try <dir-name>.eval.yaml first).

  • None → New eval. Found → Gap-fill.

Interview one question at a time and wait for each answer. Build every task from the user's answers, and write the spec only after the interview.

New eval

Elicit, one question at a time:

  1. Happy path: the most common successful use. What did the agent do, and what would confirm it worked?
  2. Edge case: a tricky-but-valid input that might trip the bare agent.
  3. Adversarial: what the skill should refuse or avoid.
  4. Neighbours: which other skills could an agent confuse with this one? For each one the user can point to (a path or a git repo), declare it in skills: and add a neighbour probe: a prompt that belongs to the neighbour, with activates: [<neighbour>].

Always propose a silence probe as well: unrelated work, activates: [].

Turn each answer into a task: a realistic prompt that never names the skill, an observable expect, an assert when the outcome is checkable, and activates: on execution tasks, naming the skill plus any declared skill it delegates to on that task, so the run can tell a description failure from a body failure. The harness is single-shot: nobody answers the agent's questions. If the skill asks before acting, the task judges that first turn: expect: the question, assert: nothing was done yet. Show the proposed YAML and confirm before writing.

Write the spec beside SKILL.md, named <dir-name>.eval.yaml, with skills: [./SKILL.md] and no engine: the backend and model are picked at run time with --model / --judge-model, so if the SKILL.md targets a non-default agent, tell the user which flag to pass. Then tell the user to commit the spec now, beside SKILL.md.

Gap-fill

Read the existing spec and report its tasks, grouped by the question each answers. Name any question with no task: execution tasks without activates:, no trigger probe, no deterministic assert:. Ask what behaviors are missing or under-tested before proposing or writing anything, even if the user only asked you to inspect it. Sharpen each gap into a task, show it, and confirm before writing it in.

Whose setup is measured

Runs load the user's own customizations by default (user skills, plugins, rules, settings and connectors; see REFERENCE.md for backend exceptions), which answers "does my skill work in my agent?". Isolate (--no-user-customizations, or user_customizations: false in the spec) when comparing backends or models, when the number leaves this machine (shared, published, compared with someone else's run), or when measuring the bare agent: each setup is different, so otherwise part of the delta is the setups. --ablate of the user's own skill needs no isolation, since both runs load the same setup.

Always tell the user which mode ran and what it loaded, from the report header's user customizations: line (absent means isolated), and relay any fix caliper compare suggests about it.

Phase 3 — First run

Validate the spec, then run at k=1 (commands in REFERENCE.md). Show the results. Fix any harness or config error (not a task failure) before moving on. If a run stops with Not logged in, ask the user whether to run the login command the error names. With their consent, run it (it opens a browser for them to finish) and rerun; never run it without consent. hermes model is an interactive picker, so ask the user to run it in their own terminal instead.

Show full SKILL.md (490 more words)Show less

Phase 4 — Check the tasks need the skill

Before anyone edits the skill, check that the tasks can tell it apart from the control: the declared neighbourhood with this skill removed. Run the control (--ablate <skill-name>, or skill:<skill-name> if an mcp: server shares the name) and a full run, both at k=3, then caliper compare them. Check activation first: if the skill never fired in the full run, both runs measure the same agent, and the fix is its description (Phase 5), not the tasks.

  • The skill fired, and the control scores about as well as the full run: the task doesn't need the skill, so iterating on the skill against it measures nothing. Sharpen the task with the user (a harder input, the specific rule the skill adds), then re-run both.
  • Full run clearly ahead: the task measures the skill. Keep it.

Keep the control's results path (caliper list <spec-name> shows which run was ablated). The skill isn't installed in that run, so editing SKILL.md can't move its number: re-diff against it instead of re-running it. Re-run it only when the tasks or the declared skills change.

Phase 5 — Diagnose and iterate

Read each failing task before suggesting a fix, and say where the fix belongs:

SignalWhere the fix belongs
Run exits 2 (backend misconfigured, unavailable model, failed hook or MCP server, or every attempt unusable)The environment or the task's hooks, not the skill. Fix and re-run
⊘ unusable attempts (infra_error, timeout, judge_error)Not the skill: rate limits, auth, or the judge. Fix and re-run
cheat outcomeThe task: it leaks its answer. Tighten the task or sandbox:
Activation fails: an expected skill didn't fireThat skill's description
Activation fails: an unexpected skill fired tooThe extra skill's description, or the overlap between the two
Activation fails because one of the user's own skills fired (skill: in the report header)Not the description alone: that skill is real competition in their setup. Decide with the user whether to sharpen the description or isolate the run
Activation passes, score lowThe skill's body
The judge's reasoning shows expect: was ambiguousThe task's grading: make the criterion observable, or add an assert:
Full run ≈ control, and the skill fired in the full runThe task (see Phase 4). If it never fired, the description

Then ask whether to iterate or finish.

  • Iterate: after the user edits their SKILL.md, re-run at k=3 and caliper compare it twice: against the previous full run (did the edit hold?) and against the kept control (does it still earn its place?). A skill that got worse can still beat the control. Diagnose again. At k=3 one attempt is a 33-point swing, so confirm a surprising win or loss at k≥5 before acting on it. Loop.
  • Done: confirm the full run beats the control on its execution tasks and holds against the previous full run, then remind the user to commit SKILL.md and the .eval.yaml together.

© edonadei, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in skills/grill-skill of edonadei/caliper.

  • SKILL.md
  • REFERENCE.md
  • grill-skill.eval.yaml

Open the folder on GitHubat commit 586b0df

Compare with similar skills

Grill Skill Eval Designer next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Grill Skill Eval Designer compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Grill Skill Eval Designer this skilledonadei/caliper208—~2.1kAutomated safety check: NotesMIT
Darwin Skill Optimizeralchaincyf/darwin-skill6.2k1 repos~4.7kAutomated safety check: PassMIT
Skill Release Gaterohitg00/ai-engineering-from-scratch66k—~1kAutomated safety check: PassMIT
Open-Science Skill Creatoraipoch/open-science5.5k—~1.7kAutomated safety check: PassApache-2.0
Skill JudgeshareAI-lab/Kode-CLI5.2k4 repos~7.5kAutomated safety check: PassApache-2.0
Skill Quality ReviewerGalaxy-Dawn/claude-scholar5.7k1 repos~3kAutomated safety check: PassMIT

Similar skills

  • Darwin Skill Optimizer

    alchaincyf/darwin-skill

    Scores SKILL.md files on a nine-dimension rubric, then improves them in a keep-or-revert loop with independent judge agents, test prompts, git history and human checkpoints.

    6.2k GitHub starsUsed in 1 repo~4.7k tokens
    Agent WorkflowsAuto-check passed
  • Skill Release Gate

    rohitg00/ai-engineering-from-scratch

    Evaluates an Agent Skill bundle before release for structure, trigger quality, artifact improvement, script correctness, safety, installed-tree integrity and host portability.

    66k GitHub stars~1k tokensUpdated today
    Agent WorkflowsAuto-check passed
  • Open-Science Skill Creator

    aipoch/open-science

    Creates, revises, evaluates and publishes skills in the Open-Science app through its native host.skills composer, with optional test prompts and benchmarks.

    5.5k GitHub stars~1.7k tokensUpdated today
    Agent WorkflowsAuto-check passed
  • Skill Judge

    shareAI-lab/Kode-CLI

    Evaluates the design quality of an agent skill against official specifications and patterns from existing examples, scoring it and suggesting improvements.

    5.2k GitHub starsUsed in 4 repos~7.5k tokens
    Agent WorkflowsAuto-check passed
  • Skill Quality Reviewer

    Galaxy-Dawn/claude-scholar

    Scores a skill across description, content organization, writing style and structure, then produces letter grades and a prioritized improvement plan.

    5.7k GitHub starsUsed in 1 repo~3k tokens
    Agent WorkflowsAuto-check passed
  • OpenCode Skill Creator

    antongulin/opencode-skill-creator

    Walks you through drafting, testing, evaluating and tuning a skill for OpenCode, from an intake interview to description optimization.

    172 GitHub stars~8.1k tokensUpdated 7 days ago
    Agent WorkflowsAuto-check passed

More from edonadei/caliper

  • Runs and interprets a skill's Caliper eval: how often it succeeds over repeated attempts, whether it triggers at all, and whether it beats the agent without it.

    208 GitHub stars~2k tokensUpdated today
    Auto-check: notes
  • Runs caliper's smoke evals against the real agent CLIs after a harness or MCP change, with a dry-run plan, failure triage and a report to attach to the PR.

    208 GitHub stars~664 tokensUpdated today
    Auto-check passed

Categories

Questions about Grill Skill Eval Designer

What does Grill Skill Eval Designer do?

Interviews you about what a skill should do, writes an eval spec from your answers, then loops through run, diagnose and improve until the skill is ready to ship. This skill uses the caliper CLI, installed with pipx as caliper-eval, to build evaluations for other skills. It frames an eval as four questions: does the agent pick the skill when it should and only then, does the skill do the job once it fires, does it beat the same agent without it, and does it stay good across edits.

When should I use Grill Skill Eval Designer?

Grill Skill Eval Designer fits situations like: designing the first eval for a skill that has none; finding the cases an existing skill eval does not cover; checking that a skill triggers on the right prompts and stays quiet otherwise; comparing a skill against the same agent without it.

How do I install Grill Skill Eval Designer in Claude Code?

Run `npx skills add edonadei/caliper --skill grill-skill -a claude-code`. Or copy the skill folder (skills/grill-skill in edonadei/caliper) into .claude/skills/grill-skill in your project. Claude Code loads it when a task matches its description.

How do I install Grill Skill Eval Designer in Codex?

Run `npx skills add edonadei/caliper --skill grill-skill -a codex`. Or copy the skill folder (skills/grill-skill in edonadei/caliper) into .agents/skills/grill-skill in your project. Codex loads it when a task matches its description.

Can I use Grill Skill Eval Designer in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add edonadei/caliper --skill grill-skill -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/grill-skill, .gemini/skills/grill-skill, .github/skills/grill-skill and .opencode/skills/grill-skill in your project.

What does Grill Skill Eval Designer need to run?

Going by SKILL.md and its folder, Grill Skill Eval Designer needs the command-line tools its instructions call (pipx). Our summary lists: The `caliper` CLI (`pipx install caliper-eval`); A SKILL.md file to build an eval for. Its frontmatter pre-approves these tools: Bash, Read, Write, Edit.

Does Grill Skill Eval Designer access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Grill Skill Eval Designer safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Grill Skill Eval Designer use?

Grill Skill Eval Designer is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Grill Skill Eval Designer use?

About 2.1k tokens (SKILL.md is roughly 8.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Grill Skill Eval Designer?

Skills that share tags, products or a category with Grill Skill Eval Designer: Darwin Skill Optimizer (alchaincyf/darwin-skill, 6.2k stars), Skill Release Gate (rohitg00/ai-engineering-from-scratch, 66k stars), Open-Science Skill Creator (aipoch/open-science, 5.5k stars) and Skill Judge (shareAI-lab/Kode-CLI, 5.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Grill Skill Eval Designer?

edonadei (a GitHub user) maintains it in edonadei/caliper, which has 208 GitHub stars. The repository holds 3 skills in this directory. The repository was last updated on October 9, 2026.

Source: edonadei/caliper on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.