Claude Code Agent Development
anthropics/claude-plugins-official
Explains how to write agents for Claude Code plugins: the markdown file with YAML frontmatter, trigger descriptions, model and color settings, and system prompt design.
Eval and improve a skill against golden cases — run the target skill blind in a fresh, context-free subagent on each example input, grade the artifact against the expected outcome, and let the gaps…
$ npx skills add dzhng/skills --skill eval-skills -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install dzhng/skills eval-skills --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/dzhng/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/authoring/eval-skills .claude/skills/eval-skills && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "eval-skills" agent skill from https://github.com/dzhng/skills/tree/main/skills/authoring/eval-skills into .claude/skills/eval-skills/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-skills", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/dzhng/skills/tree/main/skills/authoring/eval-skillsType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add dzhng/skills --skill eval-skills -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install dzhng/skills eval-skills --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/dzhng/skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/authoring/eval-skills .agents/skills/eval-skills && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "eval-skills" agent skill from https://github.com/dzhng/skills/tree/main/skills/authoring/eval-skills into .agents/skills/eval-skills/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-skills", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add dzhng/skills --skill eval-skills -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install dzhng/skills eval-skills --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/dzhng/skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/authoring/eval-skills .cursor/skills/eval-skills && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "eval-skills" agent skill from https://github.com/dzhng/skills/tree/main/skills/authoring/eval-skills into .cursor/skills/eval-skills/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-skills", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/dzhng/skills.git --path skills/authoring/eval-skills--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add dzhng/skills --skill eval-skills -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install dzhng/skills eval-skills --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/dzhng/skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/authoring/eval-skills .gemini/skills/eval-skills && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "eval-skills" agent skill from https://github.com/dzhng/skills/tree/main/skills/authoring/eval-skills into .gemini/skills/eval-skills/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-skills", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install dzhng/skills eval-skillsInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add dzhng/skills --skill eval-skills -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/dzhng/skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/authoring/eval-skills .github/skills/eval-skills && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "eval-skills" agent skill from https://github.com/dzhng/skills/tree/main/skills/authoring/eval-skills into .github/skills/eval-skills/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-skills", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add dzhng/skills --skill eval-skills -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install dzhng/skills eval-skills --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/dzhng/skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/authoring/eval-skills .opencode/skills/eval-skills && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "eval-skills" agent skill from https://github.com/dzhng/skills/tree/main/skills/authoring/eval-skills into .opencode/skills/eval-skills/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-skills", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
eval-skillsEval and improve a skill against golden cases — run the target skill blind in a fresh, context-free subagent on each example input, grade the artifact against the expected outcome, and let the gaps…
Eval Skills is an agent skill from dzhng/skills. Eval and improve a skill against golden cases — run the target skill blind in a fresh, context-free subagent on each example input, grade the artifact against the expected outcome, and let the gaps drive the edits. Use when the user wants to test/eval/improve/harden a skill, says "this skill keeps producing X / keeps missing Y", or hands a skill plus example input→expected-output pairs. Pairs with [write-skills](../write-skills/SKILL.md) (the authoring principles every fix obeys).
Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Agent Workflows, covering Subagents. The repository describes itself as: Reusable AI agent skills for software factories: explore ideas, write specs, implement, review, and run autonomous research. Works with Claude Code, Codex, and other… The licence is MIT.
7 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit d513228. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
gitFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Eval Skills loads about 1.9k tokens when it runs. Until then it costs about 124 tokens; SKILL.md has 1,180 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from dzhng/skills at commit d513228, republished under its MIT licence (© dzhng). 1,180 words, ~1,905 tokens.
.claude/skills/eval-skills/SKILL.md (or your agent's skills folder).Treat a skill like a function under test. Feed it example inputs in a clean room, check the artifacts against what good looks like, and let the failures drive the edits. The eval is only honest if the run is blind: the agent executing the skill must carry none of this conversation's context and must never see the expected output. Leak either and you are teaching to the test.
Confirm all three before spawning anything. If any is missing or unresolvable, stop and tell the user exactly which one and what a good version looks like. Do not invent cases, guess intent, or eval against a fuzzy wish.
SKILL.md. If you can't find it,
list the skills you can see and ask which one they mean.Validate inputs and surface first principles. Resolve the skill and read its first principles — what it's for and the standard it holds itself to; this is what the judge grades against, so if the skill doesn't make them clear, clarify with the user rather than inventing them. Settle the eval mode here too: judgment (a bar the judge applies) vs conformance (an exact task hit exactly) — ask the user if it's ambiguous. Then sharpen each case's bar — the outcome plus the smells, kept at the altitude the user cares about, never widened into a prescribed parts list unless the skill is conformance-style. Done when you can state the skill's first principles in a sentence and every case has a concrete input and a bar a competent judge could hold an artifact to.
Blind run, one fresh agent per case. Isolate every run so a misbehaving
skill can't touch the live checkout and each case starts clean. Prefer
capturing the artifact from the runner's final message — if the skill's
output is a plan or text, ask for it inline and nothing hits disk to leak.
When the skill must write files, give the runner a throwaway sandbox dir as
its only writable root, not a worktree of the live repo (worktree isolation
guards git state, not absolute-path or escaped writes). After every run,
sweep the live checkout (git status) and clean anything the run leaked —
isolation is best-effort, the sweep is the guarantee. Give the runner
only the input and the instruction to use the target skill — never the
bar, the smells, the other cases, or why you're asking. Done when you hold
one artifact per case, each from a context-free run, and the checkout is
clean.
Grade with a separate judge that applies judgment. Hand a fresh judge the artifact, the bar, and the skill's first principles — so it grades against the skill's own intent, not its personal taste — but never the expected output and never "make this pass." Grounded in those principles the judge is a competent practitioner: it decides whether the work clears the bar with defensible choices, and is explicitly free to fault both too-coarse and too-fine work. It must cite specific evidence for each verdict — a quote or pointer, not a number. Done when every part of the bar has a verdict grounded in the artifact.
Account for nondeterminism. Agents flicker. A single green is not proof. For any case that matters or any verdict that looks borderline, re-run the blind run 2–3× and report the pass rate. A skill that passes 1 of 3 is not fixed.
Diagnose each failure as skill-defect vs bad-case. A miss means either
the skill failed to drive the behavior (fixable here) or the bar was
wrong — it asked for something the skill should not do, can't express, or
it punished a defensible judgment call the skill was right to make (tell
the user; do not edit the skill to chase a wrong bar — that just encodes
the wrong reality). Name the defect against the write-skills failure
modes:
premature completion, vague completion criterion, missing rule, no leading
word, duplication, sediment, war story, no-op.
Revise via write-skills. Fix the named defect — and obey those authoring rules while you do it: sharpen the completion criterion before adding bulk, prefer one leading word over more sentences, add no no-ops. The failure is the spec for the edit; change only what the failure points at.
Re-eval all cases, not just the failed one. A fix can regress a case that was passing. Loop until every case clears its rate bar, or until you can show the skill structurally can't express a case — then report that instead of forcing it.
A short report: per case, pass rate and the cited gap; the defect each failure mapped to; the edits you made (or, if the user asked to approve first, the diff you propose); and the re-eval result. Make the before/after movement legible — this is the evidence the skill actually improved.
© dzhng, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/authoring/eval-skills of dzhng/skills.
Open the folder on GitHubat commit d513228
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders. This page covers the copy in dzhng/skills, which our catalogue first saw on October 7, 2026.
Eval Skills next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Eval Skills this skilldzhng/skills | 1k | — | ~1.9k | Automated safety check: Pass | MIT | |
| Claude Code Agent Developmentanthropics/claude-plugins-official | 38k | 7 repos | ~2.8k | Automated safety check: Pass | Apache-2.0 | |
| Subagent Driven DevelopmentAsvarox/allkaraoke | 261 | 37 repos | ~1.2k | Automated safety check: Pass | None | |
| Dispatching Parallel Agentsultralisp/ultralisp | 258 | 40 repos | ~1.5k | Automated safety check: Pass | None | |
| Paseo Advisor Second Opiniongetpaseo/paseo | 20k | 1 repos | ~756 | Automated safety check: Pass | Custom licence | |
| Task Observerrebelytics/one-skill-to-rule-them-all | 3.2k | 1 repos | ~12k | Automated safety check: Pass | CC-BY-4.0 |
anthropics/claude-plugins-official
Explains how to write agents for Claude Code plugins: the markdown file with YAML frontmatter, trigger descriptions, model and color settings, and system prompt design.
Asvarox/allkaraoke
A skill your agent uses when executing implementation plans with independent tasks in the current session
ultralisp/ultralisp
A skill your agent uses when facing 2+ independent tasks that can be worked on without shared state or sequential dependencies
getpaseo/paseo
Launches one separate agent through Paseo to give a second opinion on the current task, with a self-contained briefing and no permission to edit files.
rebelytics/one-skill-to-rule-them-all
Monitors task execution for skill improvement opportunities.
openobserve/openobserve
Splits a change into planner, coder and independent reviewer roles: you confirm a spec, a subagent implements it, and a separate reviewer checks each round's local WIP commit.
dzhng/skills
Compare screenshots against the intended design, distinguishing approved references from historical baselines.
dzhng/skills
Use Claude Code as an independent claude -p subagent when the user explicitly asks for Claude, wants a second-agent opinion from Claude, or asks to delegate a well-scoped task to Claude.
dzhng/skills
Refactor cleanly instead of layering sediment. An agent skill from dzhng/skills.
dzhng/skills
Create or revise agent skills. An agent skill from dzhng/skills.
dzhng/skills
Use the local Codex CLI as an independent second agent. An agent skill from dzhng/skills.
dzhng/skills
Audit or rewrite AGENTS.md so it holds only lasting principles.
Categories
Eval and improve a skill against golden cases — run the target skill blind in a fresh, context-free subagent on each example input, grade the artifact against the expected outcome, and let the gaps…. Eval Skills is an agent skill from dzhng/skills. Eval and improve a skill against golden cases — run the target skill blind in a fresh, context-free subagent on each example input, grade the artifact against the expected outcome, and let the gaps drive the edits.
Eval Skills fits situations like: the user wants to test/eval/improve/harden a skill; says this skill keeps producing X / keeps missing Y; hands a skill plus example input→expected-output pairs.
Run `npx skills add dzhng/skills --skill eval-skills -a claude-code`. Or copy the skill folder (skills/authoring/eval-skills in dzhng/skills) into .claude/skills/eval-skills in your project. Claude Code loads it when a task matches its description.
Run `npx skills add dzhng/skills --skill eval-skills -a codex`. Or copy the skill folder (skills/authoring/eval-skills in dzhng/skills) into .agents/skills/eval-skills in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add dzhng/skills --skill eval-skills -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/eval-skills, .gemini/skills/eval-skills, .github/skills/eval-skills and .opencode/skills/eval-skills in your project.
Going by SKILL.md and its folder, Eval Skills needs the command-line tools its instructions call (git).
SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Eval Skills is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.9k tokens (SKILL.md is roughly 7.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Eval Skills: Claude Code Agent Development (anthropics/claude-plugins-official, 38k stars), Subagent Driven Development (Asvarox/allkaraoke, 261 stars), Dispatching Parallel Agents (ultralisp/ultralisp, 258 stars) and Paseo Advisor Second Opinion (getpaseo/paseo, 20k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
dzhng (a GitHub user) maintains it in dzhng/skills, which has 1,020 GitHub stars. The repository holds 27 skills in this directory. The repository was last updated on October 5, 2026.
Source: dzhng/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.