Agent skill

Skill Improvement Loop

by Donchitos in Donchitos/Claude-Code-Game-Studios

Improves a single skill through a test, fix and retest loop, keeping or reverting each change based on how the static and category scores move.

MITAuto-check: notesAgent Workflows

Install Skill Improvement Loop

skills CLI
$ npx skills add Donchitos/Claude-Code-Game-Studios --skill skill-improve -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Donchitos/Claude-Code-Game-Studios skill-improve --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Donchitos/Claude-Code-Game-Studios.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/skill-improve .claude/skills/skill-improve && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
skill-improve
GitHub stars
26k
Token cost
~1.6k tokens
SKILL.md length
820 words
Files
1
Skills in repo
73
Repo updated
First seen
Licence
MIT

At a glance

Improves a single skill through a test, fix and retest loop, keeping or reverting each change based on how the static and category scores move.

  • Works in 7 steps: Parse Argument → Baseline Test → Diagnose → …
  • Fixing a skill that fails the framework's static checks
  • SKILL.md covers Phase 1: Parse Argument, Phase 2: Baseline Test, Phase 3: Diagnose and Phase 4: Propose Fix, plus 3 more sections
  • Calls bash and git

What it does

You name one skill, for example `/skill-improve tech-debt`, and the agent runs a test-fix-retest loop on it. It stops if the argument is missing or the skill file does not exist. The baseline comes from `/skill-test static`, which counts failures and warnings and records which of seven checks failed, and a category baseline is added when the skill has a category in the testing framework's catalog.

A NOT ASSESSED result, such as an unreadable file, ends the run, because zero failures from a file nobody could read is not a clean result. If both baselines are clean, the skill reports that nothing needs improving. Otherwise the agent reads the full skill file, diagnoses each failing check, applies targeted fixes, retests, and keeps or reverts each change depending on the score.

When your agent uses it

  • Fixing a skill that fails the framework's static checks
  • Raising a skill's category rubric score without rewriting it from scratch
  • Checking whether a skill already passes before spending time on edits

Example prompts

  • “/skill-improve tech-debt”
  • “Run the improvement loop on the sprint-plan skill and keep only fixes that raise its score.”
  • “Check whether the retrospective skill already passes all static and category checks.”

Requirements

  • The CCGS skill testing framework with `/skill-test` and its catalog file
  • Pre-approved tools (allowed-tools): Read, Glob, Grep, Write, Bash, Bash(bash "*/.claude/skills/skill-improve/../../hooks/yaml-helper.sh" resolve_config *)

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Parse Argument
  2. Baseline Test
  3. Diagnose
  4. Propose Fix
  5. Write and Retest
  6. Verdict
  7. Next Steps

What it can do on your machine

Read from SKILL.md and the folder at commit b21fa0f. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Glob
    • Grep
    • Write
    • Bash
    • Bash(bash "*/.claude/skills/skill-improve/../../hooks/yaml-helper.sh" resolve_config *)

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • bash
    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Skill Improvement Loop loads about 1.6k tokens when it runs. Until then it costs about 30 tokens; SKILL.md has 820 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~30
When it runs · the whole SKILL.md, loaded when a task matches
~1.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Read, Glob, Grep, Write, Bash, Bash(bash "*/.claude/skills/skill-improve/../../hooks/yaml-helper.sh"

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Donchitos/Claude-Code-Game-Studios at commit b21fa0f, republished under its MIT licence (© Donchitos). 820 words, ~1,585 tokens.

Download SKILL.mdSave it as .claude/skills/skill-improve/SKILL.md (or your agent's skills folder).
name
skill-improve
description
Improve a skill via a test-fix-retest loop — static checks, targeted fixes, keep or revert on score change.
allowed-tools
Read, Glob, Grep, Write, Bash, Bash(bash "*/.claude/skills/skill-improve/../../hooks/yaml-helper.sh" resolve_config *)
argument-hint
[skill-name]
user-invocable
true
model
sonnet

!bash "${CLAUDE_SKILL_DIR}/../../hooks/yaml-helper.sh" resolve_config --keys automation

Automation mode: Resolve modes.automation (project.local.yaml → project.yaml → default collaborative). Every AskUserQuestion call and every file write follows .claude/docs/automation-modes.md (collaborative asks always · guided major-only · autonomous logs and proceeds; automation_always_ask categories always prompt).

Skill Improve

Runs an improvement loop on a single skill: test → fix → retest → keep or revert.


Phase 1: Parse Argument

Read the skill name from the first argument. If missing, output usage and stop:

Usage: /skill-improve [skill-name]
Example: /skill-improve tech-debt

Verify .claude/skills/[name]/SKILL.md exists. If not, stop with: "Skill '[name]' not found."


Phase 2: Baseline Test

Run /skill-test static [name] and record the baseline score:

  • Count of FAILs
  • Count of WARNs
  • Which specific checks failed (Check 1–7)

Display to the user:

Static baseline:   [N] failures, [M] warnings
Failing: Check 4 (no ask-before-write), Check 5 (no handoff)

If the static result is NOT ASSESSED (the skill file could not be read or parsed), stop and report it with its reason: there is no score to improve, and 0 FAILs from a file nobody could read is not a clean result.

If baseline is 0 FAILs and 0 WARNs, note it and proceed to Phase 2b.

Phase 2b: Category Baseline

Look up the skill's category: field in CCGS Skill Testing Framework/catalog.yaml.

If no category: field is found, display: "Category: not yet assigned — skipping category checks." and skip to Phase 3.

If category is found, run /skill-test category [name] and record the category baseline:

  • Count of FAILs
  • Count of WARNs
  • Which specific category rubric metrics failed

Display to the user:

Category baseline: [N] failures, [M] warnings  ([category] rubric)

A category result of NOT ASSESSED (a rubric section the skill-test could not find or apply) is reported with its reason and is not a zero: the category dimension is then excluded from "already passes" below.

If BOTH static and category baselines are 0 FAILs and 0 WARNs — and neither is NOT ASSESSED — stop: "This skill already passes all static and category checks. No improvements needed."


Phase 3: Diagnose

Read the full skill file at .claude/skills/[name]/SKILL.md.

For each failing or warning static check, identify the exact gap:

  • Check 1 fail → which frontmatter field is missing
  • Check 2 fail → how many phases found vs. minimum required
  • Check 3 fail → no verdict keywords anywhere in the skill body
  • Check 4 fail → Write or Edit in allowed-tools but no ask-before-write language
  • Check 5 warn → no follow-up or next-step section at the end
  • Check 6 warn → context: fork set but fewer than 5 phases found
  • Check 7 warn → argument-hint is empty or doesn't match documented modes

For each failing or warning category check (if category was assigned in Phase 2b), identify the exact gap in the skill's text. For example:

  • If G2 fails (gate mode, director panel width): skill body does not size the PHASE-GATE panel from modes.workflow (PR at minimal, TD + PR at standard, all 4 at full)
  • If A2 fails (authoring, no per-section May-I-write): skill asks once at the end, not before each section write
  • If T3 fails (team, BLOCKED not surfaced): skill doesn't halt dependent work on blocked agent

Show the full combined diagnosis to the user before proposing any changes.


Show full SKILL.md (330 more words)Show less

Phase 4: Propose Fix

Write a targeted fix for each failure and warning. Show the proposed changes as clearly marked before/after blocks. Only change what is failing — do not rewrite sections that are passing.

Ask: "May I write this improved version to .claude/skills/[name]/SKILL.md?"

If the user says no, stop here.


Phase 5: Write and Retest

Record the current content of the skill file — the exact text, kept in this conversation — so Phase 6 can restore it.

Write the improved skill to .claude/skills/[name]/SKILL.md.

Re-run /skill-test static [name] and record the new static score — measured by that run, never stated from the edit. If a category was assigned, also re-run /skill-test category [name] and record the new category score.

Display the comparison:

Static:   Before [N] failures, [M] warnings  →  After [N'] failures, [M'] warnings
Category: Before [N] failures, [M] warnings  →  After [N'] failures, [M'] warnings  (if applicable)
Combined: [before] → [after] (improved / no change / worse)

[before] and [after] are the Phase 6 combined counts — failures plus warnings, both dimensions — e.g. Combined: 3 → 0 (improved).


Phase 6: Verdict

Count the combined failure total: static FAILs + category FAILs + static WARNs + category WARNs. A re-test that comes back NOT ASSESSED (for example, the edit broke the frontmatter so the file no longer parses) is never an improvement, whatever its counts: take the "did not improve" branch below.

If combined score improved (combined failure count is lower than baseline): Report: "Score improved. Changes kept." Show a summary of what was fixed in each dimension.

If combined score is the same or worse: Report: "Combined score did not improve." Show what changed and why it may not have helped. Ask: "May I restore .claude/skills/[name]/SKILL.md to the content it had before this run?" If yes: write the content recorded in Phase 5 back to the file with Write; if no, leave the file as written. Never git checkout it — that returns the last committed version and discards any edits made before this run that were not yet committed.


Phase 7: Next Steps

  • Run /skill-test static all to find the next skill with failures.
  • Run /skill-improve [next-name] to continue the loop on another skill.
  • Run /skill-test audit to see overall coverage progress.

© Donchitos, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/skill-improve of Donchitos/Claude-Code-Game-Studios.

Open the folder on GitHubat commit b21fa0f

Compare with similar skills

Skill Improvement Loop next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Skill Improvement Loop compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Skill Improvement Loop this skillDonchitos/Claude-Code-Game-Studios26k—~1.6kAutomated safety check: NotesMIT
Darwin Skill Optimizeralchaincyf/darwin-skill6.2k1 repos~4.7kAutomated safety check: PassMIT
Skill Release Gaterohitg00/ai-engineering-from-scratch65k—~1kAutomated safety check: PassMIT
Open-Science Skill Creatoraipoch/open-science5.4k—~1.7kAutomated safety check: PassApache-2.0
Skill JudgeshareAI-lab/Kode-CLI5.2k4 repos~7.5kAutomated safety check: PassApache-2.0
Skill Quality ReviewerGalaxy-Dawn/claude-scholar5.7k1 repos~3kAutomated safety check: PassMIT

Similar skills

  • Darwin Skill Optimizer

    alchaincyf/darwin-skill

    Scores SKILL.md files on a nine-dimension rubric, then improves them in a keep-or-revert loop with independent judge agents, test prompts, git history and human checkpoints.

    6.2k GitHub starsUsed in 1 repo~4.7k tokens
    Agent WorkflowsAuto-check passed
  • Skill Release Gate

    rohitg00/ai-engineering-from-scratch

    Evaluates an Agent Skill bundle before release for structure, trigger quality, artifact improvement, script correctness, safety, installed-tree integrity and host portability.

    65k GitHub stars~1k tokensUpdated today
    Agent WorkflowsAuto-check passed
  • Open-Science Skill Creator

    aipoch/open-science

    Creates, revises, evaluates and publishes skills in the Open-Science app through its native host.skills composer, with optional test prompts and benchmarks.

    5.4k GitHub stars~1.7k tokensUpdated today
    Agent WorkflowsAuto-check passed
  • Skill Judge

    shareAI-lab/Kode-CLI

    Evaluates the design quality of an agent skill against official specifications and patterns from existing examples, scoring it and suggesting improvements.

    5.2k GitHub starsUsed in 4 repos~7.5k tokens
    Agent WorkflowsAuto-check passed
  • Skill Quality Reviewer

    Galaxy-Dawn/claude-scholar

    Scores a skill across description, content organization, writing style and structure, then produces letter grades and a prioritized improvement plan.

    5.7k GitHub starsUsed in 1 repo~3k tokens
    Agent WorkflowsAuto-check passed
  • OpenCode Skill Creator

    antongulin/opencode-skill-creator

    Walks you through drafting, testing, evaluating and tuning a skill for OpenCode, from an intake interview to description optimization.

    172 GitHub stars~8.1k tokensUpdated 6 days ago
    Agent WorkflowsAuto-check passed

More from Donchitos/Claude-Code-Game-Studios

All 73 skills in this repo
  • Game Asset Audit

    Donchitos/Claude-Code-Game-Studios

    Audits game assets against naming conventions, file size budgets and format standards, and finds orphaned assets and missing references.

    26k GitHub stars~2k tokensUpdated 8 days ago
    Auto-check passed
  • Game Asset Spec Writer

    Donchitos/Claude-Code-Game-Studios

    Writes per-asset visual specs and AI image-generation prompts for a game's characters, enemies and screens, driven by the GDD, art bible and an entity inventory.

    26k GitHub stars~5k tokensUpdated 8 days ago
    Auto-check passed
  • Game Balance Check

    Donchitos/Claude-Code-Game-Studios

    Checks game data and formulas for balance outliers, broken progression, degenerate strategies and economy problems, and answers 'could not run' when the data is missing.

    26k GitHub stars~2.2k tokensUpdated 8 days ago
    Auto-check passed
  • Structured Bug Reports

    Donchitos/Claude-Code-Game-Studios

    Turns a description into a structured bug report, or scans code for likely bugs, then verifies and closes reports through four modes.

    26k GitHub stars~2.5k tokensUpdated 8 days ago
    Auto-check: notes
  • Bug Triage

    Donchitos/Claude-Code-Game-Studios

    Reviews the open bug backlog, separates severity from priority, assigns fixes to sprints and reports systemic trends, writing a dated triage file.

    26k GitHub stars~2.3k tokensUpdated 8 days ago
    Auto-check passed
  • Changelog Generator for Games

    Donchitos/Claude-Code-Game-Studios

    Generates an internal or player-facing changelog from git commits and sprint data, filtering out framework maintenance commits so that only work on the game itself reaches release copy.

    26k GitHub stars~2.5k tokensUpdated 8 days ago
    Auto-check: notes

Categories

Questions about Skill Improvement Loop

What does Skill Improvement Loop do?

Improves a single skill through a test, fix and retest loop, keeping or reverting each change based on how the static and category scores move. You name one skill, for example `/skill-improve tech-debt`, and the agent runs a test-fix-retest loop on it. It stops if the argument is missing or the skill file does not exist.

When should I use Skill Improvement Loop?

Skill Improvement Loop fits situations like: fixing a skill that fails the framework's static checks; raising a skill's category rubric score without rewriting it from scratch; checking whether a skill already passes before spending time on edits.

How do I install Skill Improvement Loop in Claude Code?

Run `npx skills add Donchitos/Claude-Code-Game-Studios --skill skill-improve -a claude-code`. Or copy the skill folder (.claude/skills/skill-improve in Donchitos/Claude-Code-Game-Studios) into .claude/skills/skill-improve in your project. Claude Code loads it when a task matches its description.

How do I install Skill Improvement Loop in Codex?

Run `npx skills add Donchitos/Claude-Code-Game-Studios --skill skill-improve -a codex`. Or copy the skill folder (.claude/skills/skill-improve in Donchitos/Claude-Code-Game-Studios) into .agents/skills/skill-improve in your project. Codex loads it when a task matches its description.

Can I use Skill Improvement Loop in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Donchitos/Claude-Code-Game-Studios --skill skill-improve -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/skill-improve, .gemini/skills/skill-improve, .github/skills/skill-improve and .opencode/skills/skill-improve in your project.

What does Skill Improvement Loop need to run?

Going by SKILL.md and its folder, Skill Improvement Loop needs the command-line tools its instructions call (bash and git). Our summary lists: The CCGS skill testing framework with `/skill-test` and its catalog file. Its frontmatter pre-approves these tools: Read, Glob, Grep, Write, Bash, Bash(bash "*/.claude/skills/skill-improve/../../hooks/yaml-helper.sh" resolve_config *).

Does Skill Improvement Loop access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Skill Improvement Loop safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Skill Improvement Loop use?

Skill Improvement Loop is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Skill Improvement Loop use?

About 1.6k tokens (SKILL.md is roughly 6.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Skill Improvement Loop?

Skills that share tags, products or a category with Skill Improvement Loop: Darwin Skill Optimizer (alchaincyf/darwin-skill, 6.2k stars), Skill Release Gate (rohitg00/ai-engineering-from-scratch, 65k stars), Open-Science Skill Creator (aipoch/open-science, 5.4k stars) and Skill Judge (shareAI-lab/Kode-CLI, 5.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Skill Improvement Loop?

Donchitos (a GitHub user) maintains it in Donchitos/Claude-Code-Game-Studios, which has 25,834 GitHub stars. The repository holds 73 skills in this directory. The repository was last updated on September 29, 2026.

Source: Donchitos/Claude-Code-Game-Studios on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.