Agent skill

Evaluate Skill with Caliper

by edonadei in edonadei/caliper

Runs and interprets a skill's Caliper eval: how often it succeeds over repeated attempts, whether it triggers at all, and whether it beats the agent without it.

MITAuto-check: notesAgent Workflows

Install Evaluate Skill with Caliper

skills CLI
$ npx skills add edonadei/caliper --skill evaluate-skill -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install edonadei/caliper evaluate-skill --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/edonadei/caliper.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/evaluate-skill .claude/skills/evaluate-skill && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
evaluate-skill
GitHub stars
206
Token cost
~1.9k tokens
SKILL.md length
1,029 words
Files
3
Skills in repo
3
Repo updated
First seen
Licence
MIT

At a glance

Runs and interprets a skill's Caliper eval: how often it succeeds over repeated attempts, whether it triggers at all, and whether it beats the agent without it.

  • Works in 4 steps: caliper validate . It never touches the… → caliper run --k 1 to shake out spec and… → caliper run --k 3 --ablate , once (write… → …
  • Running an eval for a skill you wrote and reading the results
  • SKILL.md covers Prerequisites, The four questions, Spec shape and Running, plus 4 more sections
  • Calls pipx

What it does

The skill operates the caliper command-line tool against an .eval.yaml spec and traces each failure to the place that fixes it. It separates four questions that a single score would blur: whether the skill fires when it should, whether it works once it fires, whether it earns its place against a control run without it, and whether it holds up across edits over time. Each has its own measurement and its own fix location, such as the skill's description, its body, the tasks, or the edit that moved the result.

When writing a spec, every task that checks results also checks activation, and prompts read like a real user's request with the skill left unnamed. The engine is chosen at run time and can differ for the skill and the judge, with claude-code, codex, pi and hermes available, and each attempt runs in a fresh empty working directory. Validation never touches the network, and a first run with one attempt is meant to shake out spec and harness errors before any score is trusted. A REFERENCE.md file covers the full spec format.

When your agent uses it

  • Running an eval for a skill you wrote and reading the results
  • Finding out whether a skill triggers on the prompts it should
  • Checking whether a skill beats the agent without it
  • Comparing eval runs after editing a skill
  • Writing an .eval.yaml spec once the tasks are decided

Example prompts

  • “Run the eval for my pdf-filler skill and tell me where it fails.”
  • “Validate release-notes.eval.yaml and run it with one attempt to catch harness errors.”
  • “Compare this eval run against the previous one and say which edit moved the success rate.”
  • “Run the same spec on codex and use a different model as the judge.”

Requirements

  • The caliper CLI, installed with pipx install caliper-eval
  • Bash access
  • Pre-approved tools (allowed-tools): Bash

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. caliper validate . It never touches the network.
  2. caliper run --k 1 to shake out spec and harness errors. Fix those before reading any score.
  3. caliper run --k 3 --ablate , once (write skill: if an mcp: server shares the name). This is the control: the declared neighbourhood…
  4. caliper run --k 3, then caliper compare to see whether it earns its place. After each edit, also compare against the previous full run's…

What it can do on your machine

Read from SKILL.md and the folder at commit f3f0355. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Bash

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pipx

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pipx, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Evaluate Skill with Caliper loads about 1.9k tokens when it runs. Until then it costs about 76 tokens; SKILL.md has 1,029 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~76
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Bash

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from edonadei/caliper at commit f3f0355, republished under its MIT licence (© edonadei). 1,029 words, ~1,897 tokens.

Download SKILL.mdSave it as .claude/skills/evaluate-skill/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
evaluate-skill
description
Run, read, and diagnose a skill's Caliper eval — its success rate over k attempts, whether it fires, and whether it beats the agent without it. Use when the user wants to run, validate, interpret, or compare a skill's eval, or write an .eval.yaml spec whose tasks they have already decided.
allowed-tools
Bash

Evaluate Skill

Operate Caliper: run a skill's eval, read what the results say, and trace each failure to the place that fixes it.

Prerequisites

The caliper CLI must be on PATH. This skill can be copied into an agent without the Caliper repo, so install it if missing:

bash
pipx install caliper-eval

caliper <command> --help is the authority on flags. REFERENCE.md covers the full spec format, engines, and what the results mean.

The four questions

A skill eval answers four separate questions. Each has its own measurement and its own fix location, so keep them apart: a blended number hides which half broke.

QuestionMeasured byA failure is fixed in
Fires: does the agent reach for the skill when it should, and only then?activates: on tasks, and trigger probesthe skill's description frontmatter
Works: once it fires, does it get the job done?expect: / assert:, scored as the success ratethe skill's body
Earns: does it beat the agent without it?the control (--ablate <skill-name>: the declared neighbourhood minus this skill) and caliper compare. For a truly bare agent, ablate every declared skill and every declared mcp: server, and isolate the runthe tasks: if the control passes too, the task is too easy
Holds: does it stay good across edits and over time?caliper compare of each full run against the previous one, skill driftthe edit that moved it

Spec shape

yaml
skills:
  - ./SKILL.md              # relative to the spec; installed, never preloaded
tasks:
  - name: What success looks like
    prompt: <what a real user would type; never names the skill>
    expect: <natural-language pass/fail criterion>
    assert: |               # optional deterministic Python check
      assert ...
    activates: [my-skill]   # frontmatter name: of the skill that should fire
  - name: Unrelated work stays unrelated
    prompt: <work no declared skill should answer>
    activates: []           # a trigger probe: no judge, cheap

When you write a spec, every task with expect: or assert: also asserts activates:, refusals included: the skill, plus any declared skill it delegates to on that task, and every prompt reads like a real user's request with the skill left unnamed.

The spec has no backend/model or judge: block. The engine is chosen at run time, independently for the skill and the judge: caliper run <spec> --model codex runs and grades on codex, and --judge-model picks a different judge. Backends are claude-code (default), codex, pi, and hermes. Each attempt runs in a fresh, empty workdir, so setup: builds fixtures there with relative paths.

Running

  1. caliper validate <spec>. It never touches the network.
  2. caliper run <spec> --k 1 to shake out spec and harness errors. Fix those before reading any score.
  3. caliper run <spec> --k 3 --ablate <skill-name>, once (write skill:<skill-name> if an mcp: server shares the name). This is the control: the declared neighbourhood without this skill. The skill isn't installed, so editing SKILL.md can't move its number. Keep its results path (caliper list <spec-name> marks ablated runs) and re-diff against it. Re-run it only when the tasks or the declared skills change.
  4. caliper run <spec> --k 3, then caliper compare <control.json> <spec-name> to see whether it earns its place. After each edit, also compare against the previous full run's path to see whether the edit held: a skill that got worse can still beat the control. A bare spec name resolves to that spec's latest run.
Show full SKILL.md (546 more words)Show less

Reading results

The success rate (successes / usable) is the headline. pass@k and pass^k are secondary views under --verbose. Activation is a separate scoreboard, over a different population: report it beside the success rate, never averaged into it.

Trace every failing task to where its fix belongs:

SignalWhere the fix belongs
Run exits 2 (backend misconfigured, unavailable model, failed hook or MCP server, or every attempt unusable)The environment or the task's hooks. Report it as a configuration problem, not a task failure
⊘ unusable attempts (infra_error, timeout, judge_error)Not the skill: rate limits, auth, or the judge. They are excluded from the score; re-run
cheat outcomeThe task: it leaks its answer. Tighten the task or sandbox:
Activation fails: an expected skill didn't fireThat skill's description
Activation fails: an unexpected skill fired tooThe extra skill's description, or the overlap between the two
Activation fails because one of the user's own skills fired (skill: in the report header)Not the description alone: that skill is real competition in their setup. Decide with the user whether to sharpen the description or isolate the run
Activation passes, score lowThe skill's body
The judge's reasoning shows expect: was ambiguousThe task's grading: make the criterion observable, or add an assert:
Full run ≈ control, and the skill fired in the full runThe task: it doesn't need the skill. If the skill never fired, the fix is its description (above)

At k=3 one attempt is a 33-point swing. Before calling a change a win or a regression, re-run at k≥5. caliper compare flags any drop and never gates. A rewritten or shortened skill is safe to ship when, at k≥5, it stays within about 5% of the previous score and still beats the control.

Whose setup is measured

Runs load the user's own customizations by default (user skills, plugins, rules, settings and connectors; see REFERENCE.md for backend exceptions), which answers "does my skill work in my agent?". Isolate (--no-user-customizations, or user_customizations: false in the spec) when comparing backends or models, when the number leaves this machine (shared, published, compared with someone else's run), or when measuring the bare agent: each setup is different, so otherwise part of the delta is the setups. --ablate of the user's own skill needs no isolation, since both runs load the same setup.

Always tell the user which mode ran and what it loaded, from the report header's user customizations: line (absent means isolated), and relay any fix caliper compare suggests about it.

No eval yet?

If the user hasn't decided what to test (a skill with no .eval.yaml, or an eval whose gaps they want found), suggest grill-skill: it interviews them and writes the spec. To design tasks yourself, follow "Designing evals" in REFERENCE.md.

Done when

  • every task has an observable criterion, and at least one has a deterministic assert:;
  • execution tasks assert activates:, and the spec has at least one trigger probe;
  • the full run beats the control on execution tasks, and holds against the previous full run. Trigger probes have no execution score: read them from the full run's activation;
  • the spec passes caliper validate;
  • the user has been told to commit the .eval.yaml beside SKILL.md. Saved runs under .caliper/results/ are useful for diffing over time and safe to gitignore.

© edonadei, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in skills/evaluate-skill of edonadei/caliper.

  • SKILL.md
  • REFERENCE.md
  • evaluate-skill.eval.yaml

Open the folder on GitHubat commit f3f0355

Compare with similar skills

Evaluate Skill with Caliper next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Evaluate Skill with Caliper compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Evaluate Skill with Caliper this skilledonadei/caliper206—~1.9kAutomated safety check: NotesMIT
Darwin Skill Optimizeralchaincyf/darwin-skill6.2k1 repos~4.7kAutomated safety check: PassMIT
Skill Release Gaterohitg00/ai-engineering-from-scratch65k—~1kAutomated safety check: PassMIT
Open-Science Skill Creatoraipoch/open-science5.4k—~1.7kAutomated safety check: PassApache-2.0
Skill JudgeshareAI-lab/Kode-CLI5.2k4 repos~7.5kAutomated safety check: PassApache-2.0
Skill Quality ReviewerGalaxy-Dawn/claude-scholar5.7k1 repos~3kAutomated safety check: PassMIT

Similar skills

  • Darwin Skill Optimizer

    alchaincyf/darwin-skill

    Scores SKILL.md files on a nine-dimension rubric, then improves them in a keep-or-revert loop with independent judge agents, test prompts, git history and human checkpoints.

    6.2k GitHub starsUsed in 1 repo~4.7k tokens
    Agent WorkflowsAuto-check passed
  • Skill Release Gate

    rohitg00/ai-engineering-from-scratch

    Evaluates an Agent Skill bundle before release for structure, trigger quality, artifact improvement, script correctness, safety, installed-tree integrity and host portability.

    65k GitHub stars~1k tokensUpdated today
    Agent WorkflowsAuto-check passed
  • Open-Science Skill Creator

    aipoch/open-science

    Creates, revises, evaluates and publishes skills in the Open-Science app through its native host.skills composer, with optional test prompts and benchmarks.

    5.4k GitHub stars~1.7k tokensUpdated today
    Agent WorkflowsAuto-check passed
  • Skill Judge

    shareAI-lab/Kode-CLI

    Evaluates the design quality of an agent skill against official specifications and patterns from existing examples, scoring it and suggesting improvements.

    5.2k GitHub starsUsed in 4 repos~7.5k tokens
    Agent WorkflowsAuto-check passed
  • Skill Quality Reviewer

    Galaxy-Dawn/claude-scholar

    Scores a skill across description, content organization, writing style and structure, then produces letter grades and a prioritized improvement plan.

    5.7k GitHub starsUsed in 1 repo~3k tokens
    Agent WorkflowsAuto-check passed
  • OpenCode Skill Creator

    antongulin/opencode-skill-creator

    Walks you through drafting, testing, evaluating and tuning a skill for OpenCode, from an intake interview to description optimization.

    172 GitHub stars~8.1k tokensUpdated 5 days ago
    Agent WorkflowsAuto-check passed

More from edonadei/caliper

  • Runs caliper's smoke evals against the real agent CLIs after a harness or MCP change, with a dry-run plan, failure triage and a report to attach to the PR.

    206 GitHub stars~664 tokensUpdated 2 days ago
    Auto-check passed
  • Interviews you about what a skill should do, writes an eval spec from your answers, then loops through run, diagnose and improve until the skill is ready to ship.

    206 GitHub stars~2k tokensUpdated 2 days ago
    Auto-check: notes

Categories

Questions about Evaluate Skill with Caliper

What does Evaluate Skill with Caliper do?

Runs and interprets a skill's Caliper eval: how often it succeeds over repeated attempts, whether it triggers at all, and whether it beats the agent without it. yaml spec and traces each failure to the place that fixes it. It separates four questions that a single score would blur: whether the skill fires when it should, whether it works once it fires, whether it earns its place against a control run without it, and whether it holds up across edits over time.

When should I use Evaluate Skill with Caliper?

Evaluate Skill with Caliper fits situations like: running an eval for a skill you wrote and reading the results; finding out whether a skill triggers on the prompts it should; checking whether a skill beats the agent without it; comparing eval runs after editing a skill.

How do I install Evaluate Skill with Caliper in Claude Code?

Run `npx skills add edonadei/caliper --skill evaluate-skill -a claude-code`. Or copy the skill folder (skills/evaluate-skill in edonadei/caliper) into .claude/skills/evaluate-skill in your project. Claude Code loads it when a task matches its description.

How do I install Evaluate Skill with Caliper in Codex?

Run `npx skills add edonadei/caliper --skill evaluate-skill -a codex`. Or copy the skill folder (skills/evaluate-skill in edonadei/caliper) into .agents/skills/evaluate-skill in your project. Codex loads it when a task matches its description.

Can I use Evaluate Skill with Caliper in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add edonadei/caliper --skill evaluate-skill -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/evaluate-skill, .gemini/skills/evaluate-skill, .github/skills/evaluate-skill and .opencode/skills/evaluate-skill in your project.

What does Evaluate Skill with Caliper need to run?

Going by SKILL.md and its folder, Evaluate Skill with Caliper needs the command-line tools its instructions call (pipx). Our summary lists: The caliper CLI, installed with pipx install caliper-eval; Bash access. Its frontmatter pre-approves these tools: Bash.

Does Evaluate Skill with Caliper access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Evaluate Skill with Caliper safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Evaluate Skill with Caliper use?

Evaluate Skill with Caliper is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Evaluate Skill with Caliper use?

About 1.9k tokens (SKILL.md is roughly 7.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Evaluate Skill with Caliper?

Skills that share tags, products or a category with Evaluate Skill with Caliper: Darwin Skill Optimizer (alchaincyf/darwin-skill, 6.2k stars), Skill Release Gate (rohitg00/ai-engineering-from-scratch, 65k stars), Open-Science Skill Creator (aipoch/open-science, 5.4k stars) and Skill Judge (shareAI-lab/Kode-CLI, 5.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Evaluate Skill with Caliper?

edonadei (a GitHub user) maintains it in edonadei/caliper, which has 206 GitHub stars. The repository holds 3 skills in this directory. The repository was last updated on October 5, 2026.

Source: edonadei/caliper on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.