Agent skill

Optimizing Skills

by oaustegard in oaustegard/claude-skills

Disciplined, validation-gated revision of an EXISTING skill so each edit is a measured improvement rather than a guess.

MITAuto-check passed

Install Optimizing Skills

skills CLI
$ npx skills add oaustegard/claude-skills --skill optimizing-skills -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install oaustegard/claude-skills optimizing-skills --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/oaustegard/claude-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/optimizing-skills .claude/skills/optimizing-skills && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
optimizing-skills
GitHub stars
150
Token cost
~2.3k tokens
SKILL.md length
1,190 words
Files
4 (incl. references)
Skills in repo
66
Repo updated
First seen
Licence
MIT

At a glance

Disciplined, validation-gated revision of an EXISTING skill so each edit is a measured improvement rather than a guess.

  • Works in 4 steps: Assemble a held-out check set. 3–8… → Hold two versions. best = what you… → Score both on the check set. "Run" here… → …
  • Tuning a skill that already exists and there is evidence it underperforms (observed failures
  • SKILL.md covers Core principle, The gate — run it every revision, Bounded edits (the "textual… and Reflect: failures first, then…, plus 7 more sections
  • Calls python3

What it does

Optimizing Skills is an agent skill from oaustegard/claude-skills. Disciplined, validation-gated revision of an EXISTING skill so each edit is a measured improvement rather than a guess. Use when editing, revising, or tuning a skill that already exists and there is evidence it underperforms (observed failures, drift, complaints) — invoke by name, or have versioning-skills / creating-skill defer to it before applying edits. Not for authoring a brand-new skill from scratch (use creating-skill) or one-off prose.

Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including reference files (for example `CHANGELOG.md`, `README.md` and `references/skillopt-provenance.md`).

The repository describes itself as: My collection of Claude skills. The licence is MIT.

When your agent uses it

  • Tuning a skill that already exists and there is evidence it underperforms (observed failures
  • Complaints) — invoke by name
  • Have versioning-skills / creating-skill defer to it before applying edits

Example prompts

  • “/optimizing-skills”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Assemble a held-out check set. 3–8 representative tasks/prompts the skill
  2. Hold two versions. best = what you currently ship (never let it
  3. Score both on the check set. "Run" here = dispatch each check task to the
  4. Accept only if candidate strictly beats best on the triggering-failure

What it can do on your machine

Read from SKILL.md and the folder at commit cf49d47. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Optimizing Skills loads about 2.3k tokens when it runs, and up to ~3.1k if it reads all its reference files. Until then it costs about 116 tokens; SKILL.md has 1,190 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~116
When it runs · the whole SKILL.md, loaded when a task matches
~2.3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~3.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from oaustegard/claude-skills at commit cf49d47, republished under its MIT licence (© oaustegard). 1,190 words, ~2,279 tokens.

Download SKILL.mdSave it as .claude/skills/optimizing-skills/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
optimizing-skills
description
Disciplined, validation-gated revision of an EXISTING skill so each edit is a measured improvement rather than a guess. Use when editing, revising, or tuning a skill that already exists and there is evidence it underperforms (observed failures, drift, complaints) — invoke by name, or have versioning-skills / creating-skill defer to it before applying edits. Not for authoring a brand-new skill from scratch (use creating-skill) or one-off prose.
metadata.version
0.3.0

Optimizing Skills

Treat the skill document as the parameter under optimization: change it only when the change demonstrably beats the version you already ship. This is the discipline distilled from SkillOpt (microsoft/SkillOpt, arXiv:2605.23904) — its training apparatus dropped, its reproducibility discipline kept. The point is to stop editing skills on intuition and start editing them on evidence.

Core principle

A skill edit is only worth shipping if it strictly improves measured behavior on a held-out check. Most edits that feel like improvements don't move the needle, and some quietly regress. The gate below is what separates a real improvement from a confident guess.

The gate — run it every revision

  1. Assemble a held-out check set. 3–8 representative tasks/prompts the skill should handle well, and it must include the failure(s) that prompted this revision. Keep the set fixed across the revision so before/after scores are comparable.
  2. Hold two versions. best = what you currently ship (never let it silently degrade). candidate = best + your proposed edits.
  3. Score both on the check set. "Run" here = dispatch each check task to the Agent tool (subagent_type=general-purpose) with the skill version in context, or evaluate by hand for small sets. Give the scoring agent the skill version and the task, and nothing else. Never hand it the ledger, your revision notes, or the diagnosis that motivated the edit: it will solve the task from those instead of from the skill, and the score stops measuring the skill. WikiSkill (arXiv:2608.27454) ablated exactly this and lost 2.8 points of final quality, 7.8 on their hardest split, by letting the worker read the improver's knowledge store. Score per criterion, not one collapsed pass/fail. When a task carries several criteria, the criterion that decides accept/reject is the failure that prompted this revision; the others are regression guards that must not get worse. Collapsing criteria masks the win: in the down-skilling-v1.2.0 retro, the edit drove architectural hallucination 60%→0% while an unrelated length criterion stayed 0/5 in both arms — a single combined pass/fail scored that as a 0–0 tie and would have rejected a large, real improvement.
  4. Accept only if candidate strictly beats best on the triggering-failure criterion, with no regression guard worse. Ties → reject, keep best. An edit that doesn't move the needle does not ship.

When the skill's own output is compiled by an Agent (down-skilling and creating-skill produce a prompt an author writes from the SKILL), score ≥2 author samples per version, or fix one author across both arms. A single author sample per arm lets author capability dominate the edit effect: the same down-skilling edit measured 95%→0% with one author pair and 60%→0% with another — real either way, but n=1 cannot tell a real edit from a lucky author.

This two-tier best/candidate split is the heart of it: a working revision can explore, but the shipped skill only ever ratchets upward.

Bounded edits (the "textual learning rate")

Cap edits per revision — default ~4 distinct add/replace/delete operations, fewer as the skill matures. Large speculative rewrites drift and destroy your ability to attribute a regression to a cause. If a revision wants more edits than the budget, rank and keep the top ones (below) and let the rest wait.

Reflect: failures first, then successes

Separate the evidence before proposing edits:

  • Failure reflection. Across the failing cases, find the common, systematic pattern — not a one-off edge case. Propose edits that fix the pattern. Failures take priority in any merge.
  • Success reflection. Across cases that already work, find generalizable patterns worth encoding so they survive future edits. Reinforce; don't duplicate.

For both: edits must generalize (never hardcode task-specific values), and must not duplicate content already in the skill — patch genuine gaps only.

Rank when over budget

When candidate edits exceed the budget, keep them in this priority order:

  1. Systematic impact — fixes a recurring failure across many cases, not one.
  2. Complementarity — fills a real gap rather than restating existing content.
  3. Generality — phrased as a durable principle, not tied to one task/entity.
  4. Actionability — concrete, followable guidance over vague advice.

Drop the rest. They can return next revision if still warranted.

Show full SKILL.md (509 more words)Show less

Protect the hard-won core

If a skill has a battle-tested core that routine edits keep eroding, fence it off and treat it as off-limits to fast edits. Revisit it only on a deliberate longitudinal review: compare the same check tasks across several versions to catch slow drift and regressions that single-edit review misses. (SkillOpt fences this region with HTML-comment markers and only rewrites it at epoch boundaries — the same idea, manual cadence.)

Recording rejected proposals

FIRST action of any revision, before you score anything:

bash
python3 scripts/skill_ledger.py check --skill <name> --before best.md --after candidate.md

Exit 1 means this exact edit was proposed before and failed the gate. Read the recorded criterion and score, then propose something else. Do not re-score it.

After the gate decides, record the outcome — accepted or rejected:

bash
python3 scripts/skill_ledger.py record --skill <name> \
    --before best.md --after candidate.md --outcome rejected \
    --criterion "<the triggering failure>" --score-before 0.6 --score-after 0.6

The ledger computes the diff itself and appends it. A rejected candidate is the artifact worth keeping. The edit reverts and the record stays. The next session cannot see this one, which is why the record has to live outside your context rather than in it.

This is WikiSkill's skill-impact.md (arXiv:2608.27454), whose outer loop keeps a never-rolled-back record of every proposal and its fate while the skill itself rolls back. Their proposer is told in prose not to repeat rejected approaches. check makes that mechanical instead, so it no longer depends on anyone having read the ledger. The script lives in claude-workspace at scripts/skill_ledger.py.

Carry memory across revisions

The ledger above holds the artifacts: which diffs were tried and how they scored. remember() holds the judgment the diffs cannot carry. After a revision, record what you learned about editing this skill — which kinds of edits helped, which were brittle, redundant, or harmful — via remember() tagged with the skill name. Before the next revision, recall() it. This is the compounding part: each revision starts smarter than the last, the way SkillOpt's optimizer-side meta-skill conditions its future edits.

Edit mechanics

Edits are literal string operations (the Edit tool): the target text must match exactly or the edit is a silent no-op. Keep targets unique and verbatim. Prefer append / insert-after-heading / replace-exact / delete-exact, and verify each edit landed before scoring.

When NOT to use this

  • Authoring a brand-new skill from scratch → creating-skill.
  • Tracking/rolling back versions during development → versioning-skills.
  • One-off prose with no reuse → just write it.

Checklist

  • Held-out check set assembled (includes the triggering failure), fixed for the revision
  • Edits bounded (~4 max), each generalizable and non-duplicative
  • Failure patterns addressed before success reinforcement
  • Candidate scored per-criterion against best; accept decided by the triggering-failure criterion, others as regression guards; shipped only if strictly better
  • For Agent-compiled artifacts (down-skilling, creating-skill): ≥2 author samples per version, or a fixed author across arms
  • Hard-won core left untouched unless doing a deliberate longitudinal review
  • Candidate run through skill_ledger.py check before scoring; a repeat was dropped, not re-scored
  • Scoring agent given the skill version and the task only, never the ledger or the diagnosis
  • Gate outcome recorded via skill_ledger.py record, rejections included
  • Lesson about editing this skill recorded via remember()

For the deeper "dispatch reflection/scoring to the Agent tool" recipe and the adapted reflection/ranking prompt templates, see references/skillopt-provenance.md.

© oaustegard, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (references) in optimizing-skills of oaustegard/claude-skills.

  • SKILL.md
  • CHANGELOG.md
  • README.md
  • references/skillopt-provenance.md

Open the folder on GitHubat commit cf49d47

Compare with similar skills

Optimizing Skills next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Optimizing Skills compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Optimizing Skills this skilloaustegard/claude-skills150—~2.3kAutomated safety check: PassMIT
SQL Optimizationgithub/awesome-copilot40k2 repos~2.3kAutomated safety check: PassMIT
Agent Performance Optimizerruvnet/ruflo74k2 repos~3.6kAutomated safety check: PassMIT
Database Optimizerdavila7/claude-code-templates32k8 repos~2.5kAutomated safety check: PassMIT
Prompt Optimizeraffaan-m/ECC276k2 repos~2.4kAutomated safety check: PassMIT
Cost Optimizeruvnet/ruflo74k—~997Automated safety check: NotesMIT

Similar skills

  • SQL Optimization

    github/awesome-copilot

    Official

    Universal SQL performance optimization assistant for comprehensive query tuning, indexing strategies, and database performance analysis across all SQL databases (MySQL, PostgreSQL, SQL Server…

    40k GitHub starsUsed in 2 repos~2.3k tokens
    DatabasesAuto-check passed
  • Agent skill for performance-optimizer - invoke with $agent-performance-optimizer

    74k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check passed
  • Database Optimizer

    davila7/claude-code-templates

    Expert database optimizer specializing in modern performance tuning, query optimization, and scalable architectures.

    32k GitHub starsUsed in 8 repos~2.5k tokens
    DatabasesAuto-check passed
  • Prompt Optimizer

    affaan-m/ECC

    分析原始提示,识别意图和差距,匹配ECC组件(技能/命令/代理/钩子),并输出一个可直接粘贴的优化提示。仅提供咨询角色——绝不自行执行任务。触发时机:当用户说“优化提示”、“改进我的提示”、“如何编写提示”、“帮我优化这个指令”或明确要求提高提示质量时。中文等效表达同样触发:“优化prompt”、“改进prompt”、“怎么写prompt”、“帮我优化这个指令”。不触发时机:当用户希望直接执行任…

    276k GitHub starsUsed in 2 repos~2.4k tokens
    DevelopmentAuto-check passed
  • Cost Optimize

    ruvnet/ruflo

    Analyze token usage patterns and recommend cost optimizations with estimated savings

    74k GitHub stars~997 tokensUpdated today
    AI & LLM EngineeringAuto-check: notes
  • Convert 'make it faster' requests into a bounded measured optimization loop — baseline first, generate one-hypothesis variants, benchmark each against a correctness gate, and promote the fastest…

    276k GitHub starsUsed in 1 repo~664 tokens
    Auto-check passed

More from oaustegard/claude-skills

All 66 skills in this repo
  • Vega-Lite Interactive Charts

    oaustegard/claude-skills

    Builds interactive Vega-Lite charts from uploaded data: analyzes the fields, picks five to ten fitting chart types, and produces a React artifact with the data embedded inline.

    150 GitHub stars~2.1k tokensUpdated yesterday
    Auto-check passed
  • Single-File HTML Composer

    oaustegard/claude-skills

    Builds self-contained single-file HTML pages such as reports, decks, postmortems, flowcharts and prototypes from a small spec using a bundled Python composer and templates.

    150 GitHub stars~3.2k tokensUpdated yesterday
    Auto-check passed
  • Declauding

    oaustegard/claude-skills

    Rewrites model-sounding prose into plain technical writing and checks that every claim survives, for PR text, docs, commit messages and similar drafts.

    150 GitHub stars~5.2k tokensUpdated yesterday
    Auto-check passed
  • Preact Developer

    oaustegard/claude-skills

    Guides building standards-based Preact apps with native-first choices, HTM syntax, import maps and vendored ESM, from single-file demos to larger builds.

    150 GitHub stars~4.6k tokensUpdated yesterday
    Auto-check passed
  • Bluesky Zeitgeist Sampler

    oaustegard/claude-skills

    Deprecated sampler that captures short windows of the Bluesky firehose, clusters trending terms and builds an HTML report; replaced by the browsing-bluesky skill.

    150 GitHub stars~1.4k tokensUpdated yesterday
    Auto-check passed
  • Adversarial Review Before Shipping

    oaustegard/claude-skills

    Has a fresh-context adversary attack a blog post, recommendation, analysis brief or piece of code before you ship it, using a profile suited to that kind of artifact.

    150 GitHub stars~3.4k tokensUpdated yesterday
    Auto-check passed

Questions about Optimizing Skills

What does Optimizing Skills do?

Disciplined, validation-gated revision of an EXISTING skill so each edit is a measured improvement rather than a guess. Optimizing Skills is an agent skill from oaustegard/claude-skills. Disciplined, validation-gated revision of an EXISTING skill so each edit is a measured improvement rather than a guess.

When should I use Optimizing Skills?

Optimizing Skills fits situations like: tuning a skill that already exists and there is evidence it underperforms (observed failures; complaints) — invoke by name; have versioning-skills / creating-skill defer to it before applying edits.

How do I install Optimizing Skills in Claude Code?

Run `npx skills add oaustegard/claude-skills --skill optimizing-skills -a claude-code`. Or copy the skill folder (optimizing-skills in oaustegard/claude-skills) into .claude/skills/optimizing-skills in your project. Claude Code loads it when a task matches its description.

How do I install Optimizing Skills in Codex?

Run `npx skills add oaustegard/claude-skills --skill optimizing-skills -a codex`. Or copy the skill folder (optimizing-skills in oaustegard/claude-skills) into .agents/skills/optimizing-skills in your project. Codex loads it when a task matches its description.

Can I use Optimizing Skills in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add oaustegard/claude-skills --skill optimizing-skills -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/optimizing-skills, .gemini/skills/optimizing-skills, .github/skills/optimizing-skills and .opencode/skills/optimizing-skills in your project.

What does Optimizing Skills need to run?

Going by SKILL.md and its folder, Optimizing Skills needs the command-line tools its instructions call (python3). Our summary lists: Python 3.

Does Optimizing Skills access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Optimizing Skills safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Optimizing Skills use?

Optimizing Skills is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Optimizing Skills use?

About 2.3k tokens (SKILL.md is roughly 9.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 812 tokens, read only when the agent opens those files.

What are the alternatives to Optimizing Skills?

Skills that share tags, products or a category with Optimizing Skills: SQL Optimization (github/awesome-copilot, 40k stars), Agent Performance Optimizer (ruvnet/ruflo, 74k stars), Database Optimizer (davila7/claude-code-templates, 32k stars) and Prompt Optimizer (affaan-m/ECC, 276k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Optimizing Skills?

oaustegard (a GitHub user) maintains it in oaustegard/claude-skills, which has 150 GitHub stars. The repository holds 66 skills in this directory. The repository was last updated on October 8, 2026.

Source: oaustegard/claude-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.