Agent skill

Eval Skills

by dzhng in dzhng/skills

Eval and improve a skill against golden cases — run the target skill blind in a fresh, context-free subagent on each example input, grade the artifact against the expected outcome, and let the gaps…

MITAuto-check passedAgent Workflows

Install Eval Skills

skills CLI
$ npx skills add dzhng/skills --skill eval-skills -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install dzhng/skills eval-skills --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/dzhng/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/authoring/eval-skills .claude/skills/eval-skills && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
eval-skills
GitHub stars
1k
Token cost
~1.9k tokens
SKILL.md length
1,180 words
Files
1
Skills in repo
27
Repo updated
First seen
Licence
MIT

At a glance

Eval and improve a skill against golden cases — run the target skill blind in a fresh, context-free subagent on each example input, grade the artifact against the expected outcome, and let the gaps…

  • Works in 7 steps: Validate inputs and surface first… → Blind run, one fresh agent per case.… → Grade with a separate judge that applies… → …
  • The user wants to test/eval/improve/harden a skill
  • SKILL.md covers Inputs you need — refuse…, Workflow, Output and Rules
  • Calls git

What it does

Eval Skills is an agent skill from dzhng/skills. Eval and improve a skill against golden cases — run the target skill blind in a fresh, context-free subagent on each example input, grade the artifact against the expected outcome, and let the gaps drive the edits. Use when the user wants to test/eval/improve/harden a skill, says "this skill keeps producing X / keeps missing Y", or hands a skill plus example input→expected-output pairs. Pairs with [write-skills](../write-skills/SKILL.md) (the authoring principles every fix obeys).

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Agent Workflows, covering Subagents. The repository describes itself as: Reusable AI agent skills for software factories: explore ideas, write specs, implement, review, and run autonomous research. Works with Claude Code, Codex, and other… The licence is MIT.

When your agent uses it

  • The user wants to test/eval/improve/harden a skill
  • Says this skill keeps producing X / keeps missing Y
  • Hands a skill plus example input→expected-output pairs

Example prompts

  • “this skill keeps producing X / keeps missing Y”
  • “/eval-skills”

Workflow steps

7 steps, taken from the first numbered list in SKILL.md.

  1. Validate inputs and surface first principles. Resolve the skill and
  2. Blind run, one fresh agent per case. Isolate every run so a misbehaving
  3. Grade with a separate judge that applies judgment. Hand a fresh judge
  4. Account for nondeterminism. Agents flicker. A single green is not
  5. Diagnose each failure as skill-defect vs bad-case. A miss means either
  6. Revise via write-skills. Fix the named defect — and obey those
  7. Re-eval all cases, not just the failed one. A fix can regress a case

What it can do on your machine

Read from SKILL.md and the folder at commit d513228. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Eval Skills loads about 1.9k tokens when it runs. Until then it costs about 124 tokens; SKILL.md has 1,180 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~124
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from dzhng/skills at commit d513228, republished under its MIT licence (© dzhng). 1,180 words, ~1,905 tokens.

Download SKILL.mdSave it as .claude/skills/eval-skills/SKILL.md (or your agent's skills folder).
name
eval-skills
description
Eval and improve a skill against golden cases — run the target skill blind in a fresh, context-free subagent on each example input, grade the artifact against the expected outcome, and let the gaps drive the edits. Use when the user wants to test/eval/improve/harden a skill, says "this skill keeps producing X / keeps missing Y", or hands a skill plus example input→expected-output pairs. Pairs with [write-skills](../write-skills/SKILL.md) (the authoring principles every fix obeys).

Eval Skills

Treat a skill like a function under test. Feed it example inputs in a clean room, check the artifacts against what good looks like, and let the failures drive the edits. The eval is only honest if the run is blind: the agent executing the skill must carry none of this conversation's context and must never see the expected output. Leak either and you are teaching to the test.

Inputs you need — refuse without them

Confirm all three before spawning anything. If any is missing or unresolvable, stop and tell the user exactly which one and what a good version looks like. Do not invent cases, guess intent, or eval against a fuzzy wish.

  • Target skill — must resolve to a real SKILL.md. If you can't find it, list the skills you can see and ask which one they mean.
  • At least one golden case — a concrete input the skill will actually receive: a screenshot, a prompt, a file, a scene. "Improve write-spec" with no input attached is not a case.
  • The bar per case — the outcome a good artifact achieves and the smells that would make it bad, not an exhaustive parts list. The skill's judgment is what's under test, so do not pre-enumerate every requirement — that turns the eval into a conformance check and stops testing whether the skill decides well. "Sliced so each piece is independently buildable and verifiable, at the granularity a competent practitioner would pick — a lazy mega-slice and pointless over-splitting are both failures" is a bar a judge can hold the work to; "slices it well" is too thin to grade and a fixed list of expected slices is too prescriptive. State the bar and the smells; let the judge apply them. The exception is a conformance-style skill that genuinely wants an exact task hit exactly — then the explicit criteria are the bar; match the bar's shape to the skill's nature, and if you can't tell which it is, ask. If the user gives only a fuzzy wish with no bar, draw the bar out of them and echo it back before spending agents.

Workflow

  1. Validate inputs and surface first principles. Resolve the skill and read its first principles — what it's for and the standard it holds itself to; this is what the judge grades against, so if the skill doesn't make them clear, clarify with the user rather than inventing them. Settle the eval mode here too: judgment (a bar the judge applies) vs conformance (an exact task hit exactly) — ask the user if it's ambiguous. Then sharpen each case's bar — the outcome plus the smells, kept at the altitude the user cares about, never widened into a prescribed parts list unless the skill is conformance-style. Done when you can state the skill's first principles in a sentence and every case has a concrete input and a bar a competent judge could hold an artifact to.

  2. Blind run, one fresh agent per case. Isolate every run so a misbehaving skill can't touch the live checkout and each case starts clean. Prefer capturing the artifact from the runner's final message — if the skill's output is a plan or text, ask for it inline and nothing hits disk to leak. When the skill must write files, give the runner a throwaway sandbox dir as its only writable root, not a worktree of the live repo (worktree isolation guards git state, not absolute-path or escaped writes). After every run, sweep the live checkout (git status) and clean anything the run leaked — isolation is best-effort, the sweep is the guarantee. Give the runner only the input and the instruction to use the target skill — never the bar, the smells, the other cases, or why you're asking. Done when you hold one artifact per case, each from a context-free run, and the checkout is clean.

  3. Grade with a separate judge that applies judgment. Hand a fresh judge the artifact, the bar, and the skill's first principles — so it grades against the skill's own intent, not its personal taste — but never the expected output and never "make this pass." Grounded in those principles the judge is a competent practitioner: it decides whether the work clears the bar with defensible choices, and is explicitly free to fault both too-coarse and too-fine work. It must cite specific evidence for each verdict — a quote or pointer, not a number. Done when every part of the bar has a verdict grounded in the artifact.

  4. Account for nondeterminism. Agents flicker. A single green is not proof. For any case that matters or any verdict that looks borderline, re-run the blind run 2–3× and report the pass rate. A skill that passes 1 of 3 is not fixed.

  5. Diagnose each failure as skill-defect vs bad-case. A miss means either the skill failed to drive the behavior (fixable here) or the bar was wrong — it asked for something the skill should not do, can't express, or it punished a defensible judgment call the skill was right to make (tell the user; do not edit the skill to chase a wrong bar — that just encodes the wrong reality). Name the defect against the write-skills failure modes: premature completion, vague completion criterion, missing rule, no leading word, duplication, sediment, war story, no-op.

  6. Revise via write-skills. Fix the named defect — and obey those authoring rules while you do it: sharpen the completion criterion before adding bulk, prefer one leading word over more sentences, add no no-ops. The failure is the spec for the edit; change only what the failure points at.

  7. Re-eval all cases, not just the failed one. A fix can regress a case that was passing. Loop until every case clears its rate bar, or until you can show the skill structurally can't express a case — then report that instead of forcing it.

Show full SKILL.md (206 more words)Show less

Output

A short report: per case, pass rate and the cited gap; the defect each failure mapped to; the edits you made (or, if the user asked to approve first, the diff you propose); and the re-eval result. Make the before/after movement legible — this is the evidence the skill actually improved.

Rules

  • Blind is non-negotiable. The runner sees input only. The judge sees artifact + bar + the skill's first principles. The moment either sees the expected output, the eval is worthless.
  • Test judgment, not conformance. The bar is a standard the work must clear, never a checklist of the answer. If you find yourself listing the exact pieces you expect, you've stopped evaluating the skill.
  • Isolation is best-effort; the sweep is the guarantee. Always check the live checkout after a run and clean leaks, no matter how the run was sandboxed.
  • One fresh agent per case per run — no shared context, so no cross-case learning inflates a later case.
  • Grade against the bar, not against the other artifacts, and not on a numeric score that hides which part of the bar failed.
  • Don't bend the skill to pass a case you can't defend. A failing case that exposes a bad bar is a finding, not a bug.

© dzhng, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/authoring/eval-skills of dzhng/skills.

Open the folder on GitHubat commit d513228

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders. This page covers the copy in dzhng/skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Eval Skills next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Eval Skills compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Eval Skills this skilldzhng/skills1k—~1.9kAutomated safety check: PassMIT
Claude Code Agent Developmentanthropics/claude-plugins-official38k7 repos~2.8kAutomated safety check: PassApache-2.0
Subagent Driven DevelopmentAsvarox/allkaraoke26137 repos~1.2kAutomated safety check: PassNone
Dispatching Parallel Agentsultralisp/ultralisp25840 repos~1.5kAutomated safety check: PassNone
Paseo Advisor Second Opiniongetpaseo/paseo20k1 repos~756Automated safety check: PassCustom licence
Task Observerrebelytics/one-skill-to-rule-them-all3.2k1 repos~12kAutomated safety check: PassCC-BY-4.0

Similar skills

  • Claude Code Agent Development

    anthropics/claude-plugins-official

    Official

    Explains how to write agents for Claude Code plugins: the markdown file with YAML frontmatter, trigger descriptions, model and color settings, and system prompt design.

    38k GitHub starsUsed in 7 repos~2.8k tokens
    Agent WorkflowsAuto-check passed
  • Subagent Driven Development

    Asvarox/allkaraoke

    A skill your agent uses when executing implementation plans with independent tasks in the current session

    261 GitHub starsUsed in 37 repos~1.2k tokens
    Agent WorkflowsAuto-check passed
  • Dispatching Parallel Agents

    ultralisp/ultralisp

    A skill your agent uses when facing 2+ independent tasks that can be worked on without shared state or sequential dependencies

    258 GitHub starsUsed in 40 repos~1.5k tokens
    Agent WorkflowsAuto-check passed
  • Launches one separate agent through Paseo to give a second opinion on the current task, with a self-contained briefing and no permission to edit files.

    20k GitHub starsUsed in 1 repo~756 tokens
    Agent WorkflowsAuto-check passed
  • Task Observer

    rebelytics/one-skill-to-rule-them-all

    Monitors task execution for skill improvement opportunities.

    3.2k GitHub starsUsed in 1 repo~12k tokens
    Agent WorkflowsAuto-check passed
  • O2 Review Loop

    openobserve/openobserve

    Splits a change into planner, coder and independent reviewer roles: you confirm a spec, a subagent implements it, and a separate reviewer checks each round's local WIP commit.

    22k GitHub stars~3.7k tokensUpdated today
    Agent WorkflowsAuto-check passed

More from dzhng/skills

All 27 skills in this repo
  • Compare screenshots against the intended design, distinguishing approved references from historical baselines.

    1k GitHub stars~2.6k tokensUpdated 4 days ago
    Auto-check passed
  • Claude

    dzhng/skills

    Use Claude Code as an independent claude -p subagent when the user explicitly asks for Claude, wants a second-agent opinion from Claude, or asks to delegate a well-scoped task to Claude.

    1k GitHub stars~1.3k tokensUpdated 4 days ago
    Auto-check passed
  • Refactor Clean

    dzhng/skills

    Refactor cleanly instead of layering sediment. An agent skill from dzhng/skills.

    1k GitHub stars~3.1k tokensUpdated 4 days ago
    Auto-check passed
  • Write Skills

    dzhng/skills

    Create or revise agent skills. An agent skill from dzhng/skills.

    1k GitHub stars~2.7k tokensUpdated 4 days ago
    Auto-check passed
  • Codex

    dzhng/skills

    Use the local Codex CLI as an independent second agent. An agent skill from dzhng/skills.

    1k GitHub stars~2.4k tokensUpdated 4 days ago
    Auto-check: warnings
  • Audit Agents

    dzhng/skills

    Audit or rewrite AGENTS.md so it holds only lasting principles.

    1k GitHub stars~847 tokensUpdated 4 days ago
    Auto-check passed

Categories

Questions about Eval Skills

What does Eval Skills do?

Eval and improve a skill against golden cases — run the target skill blind in a fresh, context-free subagent on each example input, grade the artifact against the expected outcome, and let the gaps…. Eval Skills is an agent skill from dzhng/skills. Eval and improve a skill against golden cases — run the target skill blind in a fresh, context-free subagent on each example input, grade the artifact against the expected outcome, and let the gaps drive the edits.

When should I use Eval Skills?

Eval Skills fits situations like: the user wants to test/eval/improve/harden a skill; says this skill keeps producing X / keeps missing Y; hands a skill plus example input→expected-output pairs.

How do I install Eval Skills in Claude Code?

Run `npx skills add dzhng/skills --skill eval-skills -a claude-code`. Or copy the skill folder (skills/authoring/eval-skills in dzhng/skills) into .claude/skills/eval-skills in your project. Claude Code loads it when a task matches its description.

How do I install Eval Skills in Codex?

Run `npx skills add dzhng/skills --skill eval-skills -a codex`. Or copy the skill folder (skills/authoring/eval-skills in dzhng/skills) into .agents/skills/eval-skills in your project. Codex loads it when a task matches its description.

Can I use Eval Skills in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add dzhng/skills --skill eval-skills -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/eval-skills, .gemini/skills/eval-skills, .github/skills/eval-skills and .opencode/skills/eval-skills in your project.

What does Eval Skills need to run?

Going by SKILL.md and its folder, Eval Skills needs the command-line tools its instructions call (git).

Does Eval Skills access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Eval Skills safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Eval Skills use?

Eval Skills is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Eval Skills use?

About 1.9k tokens (SKILL.md is roughly 7.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Eval Skills?

Skills that share tags, products or a category with Eval Skills: Claude Code Agent Development (anthropics/claude-plugins-official, 38k stars), Subagent Driven Development (Asvarox/allkaraoke, 261 stars), Dispatching Parallel Agents (ultralisp/ultralisp, 258 stars) and Paseo Advisor Second Opinion (getpaseo/paseo, 20k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Eval Skills?

dzhng (a GitHub user) maintains it in dzhng/skills, which has 1,020 GitHub stars. The repository holds 27 skills in this directory. The repository was last updated on October 5, 2026.

Source: dzhng/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.