Agent skill

Skill Creator

by curiositech in curiositech/some_claude_skills

Guide for creating, improving, benchmarking, and packaging Claude Agent Skills (SKILL.md files).

Apache-2.0Auto-check passedAgent Workflows

Install Skill Creator

skills CLI
$ npx skills add curiositech/some_claude_skills --skill skill-creator -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install curiositech/some_claude_skills skill-creator --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/curiositech/some_claude_skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/corpus/meta-skills-experiment/self-improved/skill-creator .claude/skills/skill-creator && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
skill-creator
GitHub stars
243
Token cost
~5.7k tokens
SKILL.md length
2,481 words
Files
22 (incl. scripts, references, assets)
Skills in repo
109
Repo updated
First seen
Licence
Apache-2.0

At a glance

Guide for creating, improving, benchmarking, and packaging Claude Agent Skills (SKILL.md files).

  • Works in 9 steps: Spawn all runs (with-skill AND baseline)… → While runs are in progress, draft… → As runs complete, capture timing data → …
  • S want to create a skill from scratch
  • SKILL.md covers Available Resources, Communicating with the user, Creating a skill and Running and evaluating test…, plus 6 more sections
  • Runs Python scripts from its folder; calls python and claude

What it does

Skill Creator is an agent skill from curiositech/some_claude_skills. Guide for creating, improving, benchmarking, and packaging Claude Agent Skills (SKILL.md files). Invoke when users want to create a skill from scratch, improve or test an existing skill, benchmark skill performance with variance analysis, or optimize a skill description for triggering accuracy. Also invoke when users say "turn this into a skill", "make a skill for X", "help me write a SKILL.md", "my skill isn't firing correctly", or want to convert a workflow/conversation into a reusable skill. Invoke proactively…

Its SKILL.md is about 5.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 26 other files, including scripts, reference files and assets (for example `EVALUATION.md`, `PROVENANCE.md` and `agents/analyzer.md`).

It sits in Agent Workflows, covering Skill authoring. The repository describes itself as: Claude skills that make my life easier. The licence is Apache-2.0.

When your agent uses it

  • S want to create a skill from scratch
  • Test an existing skill
  • Benchmark skill performance with variance analysis
  • Optimize a skill description for triggering accuracy

Example prompts

  • “turn this into a skill”
  • “make a skill for X”
  • “help me write a SKILL.md”
  • “/skill-creator”

Requirements

  • Python 3

Workflow steps

9 steps, taken from the step headings in SKILL.md.

  1. Spawn all runs (with-skill AND baseline) in the same turn
  2. While runs are in progress, draft assertions
  3. As runs complete, capture timing data
  4. Grade, aggregate, and launch the viewer
  5. Read the feedback
  6. Generate trigger eval queries
  7. Review with user
  8. Run the optimization loop
  9. Apply the result

What it can do on your machine

Read from SKILL.md and the folder at commit 6713fc7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 3 files in scripts/ (Python, from the files we listed), which the agent can run.

    Shell commands in SKILL.md call:

    • python
    • claude

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Skill Creator loads about 5.7k tokens when it runs, and up to ~12k if it reads all its reference files. Until then it costs about 177 tokens; SKILL.md has 2,481 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~177
When it runs · the whole SKILL.md, loaded when a task matches
~5.7k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~12k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from curiositech/some_claude_skills at commit 6713fc7, republished under its Apache-2.0 licence (© curiositech). 2,481 words, ~5,747 tokens.

Download SKILL.mdSave it as .claude/skills/skill-creator/SKILL.md (or your agent's skills folder). This skill also uses 21 other files; get the full folder from GitHub.
name
skill-creator
description
Guide for creating, improving, benchmarking, and packaging Claude Agent Skills (SKILL.md files). Invoke when users want to create a skill from scratch, improve or test an existing skill, benchmark skill performance with variance analysis, or optimize a skill description for triggering accuracy. Also invoke when users say "turn this into a skill", "make a skill for X", "help me write a SKILL.md", "my skill isn't firing correctly", or want to convert a workflow/conversation into a reusable skill. Invoke proactively when a conversation has produced a repeatable workflow worth capturing. If the user mentions SKILL.md, skill files, skill descriptions, or skill triggering, this skill applies.

Skill Creator

Create new skills and iteratively improve them. The lifecycle:

  1. Decide what the skill should do and how
  2. Draft the skill
  3. Run test prompts through claude-with-the-skill
  4. Evaluate results qualitatively and quantitatively with the user
  5. Rewrite based on feedback
  6. Repeat until satisfied, then expand the test set
  7. Package and deliver

Determine where the user is in this lifecycle and help them progress. Starting from scratch ("I want a skill for X") -- help narrow scope, draft, test, iterate. Already have a draft -- go straight to eval/iterate. Just want a better description -- jump to Description Optimization.

Be flexible: if the user says "just vibe with me, no evals," do that.


Available Resources

Before starting, verify what is available in the environment. The full workflow requires these bundled resources:

ResourcePathUsed for
Grader instructionsagents/grader.mdAssertion evaluation
Comparator instructionsagents/comparator.mdBlind A/B comparison
Analyzer instructionsagents/analyzer.mdBenchmark pattern analysis
Schema referencereferences/schemas.mdevals.json, grading.json formats
Assertion patternsreferences/eval-patterns.mdWriting discriminating assertions
Troubleshootingreferences/troubleshooting.mdCommon errors and fixes
Eval review templateassets/eval_review.htmlDescription optimization UI
Benchmark aggregatorscripts/aggregate_benchmarkpython -m scripts.aggregate_benchmark
Description optimizerscripts/run_looppython -m scripts.run_loop
Eval viewereval-viewer/generate_review.pyHuman review interface

If resources are missing: The core loop (draft -> test -> review -> improve) still works without them -- grade inline and present results in-conversation instead of the browser viewer. Note what is unavailable to the user and adapt.


Communicating with the user

The skill creator serves people across a wide range of technical backgrounds. Non-developers are opening terminals for the first time because Claude makes things possible; experienced engineers are also common.

Pay attention to context cues:

  • "evaluation" and "benchmark" are borderline but generally OK
  • Avoid "JSON" or "assertion" without a brief explanation unless the user has already used those terms

Err toward clarity. A short inline definition never hurts.


Creating a skill

Capture Intent

The current conversation may already contain the workflow to capture. If the user says "turn this into a skill" or "make this repeatable," mine the conversation history first:

  • What tools were used, in what sequence?
  • What corrections did the user make?
  • What were the input and output formats?
  • What would a different user need to know to reproduce this?

Fill gaps with the user and confirm before drafting.

If starting from scratch, establish:

  1. What should this skill enable Claude to do?
  2. When should this skill trigger? (what user phrases, contexts, task types)
  3. What is the expected output -- files, structured data, prose, or a workflow?
  4. Should test cases be set up? Skills with verifiable outputs (file transforms, code generation, fixed workflows) benefit from them. Skills with subjective outputs (writing style, creative direction) often do not. Suggest the appropriate default based on skill type, but let the user decide.
Interview and Research

Ask about edge cases, input/output examples, success criteria, and dependencies before writing test prompts. If useful MCPs are available (docs search, similar skill lookup), research in parallel via subagents. Come prepared so the interview is lightweight.

Write the SKILL.md

Based on the interview, compose the skill. Every skill has:

  • name: Identifier (kebab-case)
  • description: The primary triggering mechanism. Describe what the skill does AND when to use it. All "when to trigger" information goes here, not in the body. Apply the "pushy principle" -- err toward including trigger contexts rather than omitting them. Example: instead of "Builds REST APIs," write "Builds REST APIs. Invoke whenever users ask about endpoints, routes, HTTP methods, API design, or want to add a backend -- even if they don't say 'REST.'"
  • compatibility: Required tools or dependencies (optional, rarely needed)
  • skill body: Instructions, workflow, examples, references
Skill output template

Use this as the starting structure -- fill in or remove sections as needed:

markdown
---
name: skill-name
description: >
  [What this skill does and when to trigger it. One to three sentences.
  Apply the pushy principle: include adjacent trigger contexts.]
---

# [Skill Name]

[One paragraph: what this skill enables and when to use it.]

## [Core workflow or main section]

[Instructions in imperative form. Explain the *why* behind each step.]

## Output format

[Concrete template for what this skill produces.]

## Examples

**Example 1:**
Input: [realistic user prompt]
Output: [expected result]

## Reference files

- Read `references/advanced.md` when [specific situation]
- Run `scripts/transform.py` for [repeatable operation]
Anatomy of a skill directory
skill-name/
+-- SKILL.md (required)
+-- Optional bundled resources:
    +-- scripts/    - Scripts for deterministic/repetitive tasks
    +-- references/ - Docs loaded into context as needed
    +-- assets/     - Templates, icons, fonts

Progressive loading:

  1. Metadata (name + description) -- always in context
  2. SKILL.md body -- in context when the skill triggers (<500 lines ideal)
  3. Bundled resources -- loaded as needed, scripts can run without loading

Keep SKILL.md under 500 lines. If approaching that limit, add a layer of hierarchy with clear pointers to reference files. For reference files over 300 lines, add a table of contents.

Domain organization -- when a skill covers multiple frameworks/variants:

cloud-deploy/
+-- SKILL.md (workflow + selection logic)
+-- references/
    +-- aws.md
    +-- gcp.md
    +-- azure.md

Claude reads only the relevant reference file.

Writing patterns

Output templates -- be concrete:

markdown
## Report structure
ALWAYS use this template:
# [Title]
## Executive summary
## Key findings
## Recommendations

Examples -- include realistic ones:

markdown
## Commit message format
**Example:**
Input: Added user authentication with JWT tokens
Output: feat(auth): implement JWT-based authentication

Writing style -- use imperative form. Explain the why rather than issuing rules. Avoid ALWAYS/NEVER in all caps where the reasoning can be explained instead -- models respond better to understanding than mandates. Write a draft, then read it fresh and improve it.

Security -- skills must not contain malware, exploit code, or anything that would surprise the user if they read the description. Do not create skills designed to facilitate unauthorized access or data exfiltration. Roleplay skills are fine.

Test Cases

After drafting, write 2-3 realistic test prompts -- things a real user would actually type. Share them: "Here are a few test cases I'd like to try. Do these look right, or do you want to add more?"

Save to evals/evals.json. Do not write assertions yet -- just prompts. Assertions get drafted while runs are in progress.

json
{
  "skill_name": "example-skill",
  "evals": [
    {
      "id": 1,
      "prompt": "User's task prompt",
      "expected_output": "Description of expected result",
      "files": []
    }
  ]
}

See references/schemas.md for the full schema including the assertions field.


Running and evaluating test cases

This is one continuous sequence -- do not stop partway through. Do NOT use /skill-test or any other testing skill.

Put results in <skill-name>-workspace/ as a sibling to the skill directory. Organize by iteration, then by eval name, then by configuration, then by run. Create directories as you go, not all upfront.

Workspace layout
<skill-name>-workspace/
+-- iteration-1/
|   +-- eval-<name>/                    <-- one per test case (descriptive name)
|   |   +-- eval_metadata.json          <-- eval ID, prompt, assertions
|   |   +-- with_skill/
|   |   |   +-- run-1/
|   |   |       +-- outputs/            <-- files the subagent produced
|   |   |       |   +-- metrics.json
|   |   |       +-- timing.json         <-- from task notification
|   |   |       +-- grading.json        <-- written by grader agent
|   |   +-- without_skill/              <-- or old_skill/ when improving
|   |       +-- run-1/
|   |           +-- outputs/
|   |           +-- timing.json
|   |           +-- grading.json
|   +-- benchmark.json                  <-- written by aggregate_benchmark.py
|   +-- benchmark.md
+-- iteration-2/
    +-- ...
Step 1: Spawn all runs (with-skill AND baseline) in the same turn

For each test case, spawn two subagents simultaneously -- one with the skill, one without. Do not do with-skill runs first and baseline runs later. Launch everything at once so results arrive around the same time.

With-skill run:

Execute this task:
- Skill path: <path-to-skill>
- Task: <eval prompt>
- Input files: <eval files if any, or "none">
- Save outputs to: <workspace>/iteration-<N>/eval-<name>/with_skill/run-1/outputs/
- Outputs to save: <what the user cares about>

Baseline run -- depends on context:

  • New skill: no skill at all. Same prompt, save to <workspace>/iteration-<N>/eval-<name>/without_skill/run-1/outputs/
  • Improving existing skill: snapshot the old version first (cp -r <skill-path> <workspace>/skill-snapshot/), point baseline at snapshot, save to old_skill/run-1/outputs/

Write eval_metadata.json at the eval level (e.g., <workspace>/iteration-N/eval-<name>/eval_metadata.json). Give each eval a descriptive name based on what it tests -- not just "eval-0":

json
{
  "eval_id": 0,
  "eval_name": "descriptive-name-here",
  "prompt": "The user's task prompt",
  "assertions": []
}
Step 2: While runs are in progress, draft assertions

Do not just wait for runs to finish -- use this time productively. Draft quantitative assertions for each test case and explain them to the user. If assertions already exist in evals/evals.json, review and explain them.

Good assertions are objectively verifiable and have descriptive names -- someone glancing at benchmark results should immediately understand what each one checks. For subjective skills (writing style, design quality), skip assertions and rely on qualitative review. See references/eval-patterns.md for assertion patterns by task type and the discriminating assertion test.

Update eval_metadata.json files and evals/evals.json with assertions once drafted. Preview for the user: explain what they will see in the viewer.

Step 3: As runs complete, capture timing data

When each subagent task completes, save timing data immediately to timing.json in the run directory (e.g., with_skill/run-1/timing.json):

json
{
  "total_tokens": 84852,
  "duration_ms": 23332,
  "total_duration_seconds": 23.3
}

This data only comes through task completion notifications -- capture it then.

Step 4: Grade, aggregate, and launch the viewer

GENERATE THE EVAL VIEWER BEFORE EVALUATING INPUTS YOURSELF. Get outputs in front of the human first.

Once all runs complete:

  1. Grade each run -- spawn a grader subagent for each run directory (both with_skill and without_skill/old_skill get their own grader). Use this prompt:

    You are a grader agent. Read agents/grader.md from <skill-creator-path>
    to load your full instructions, then grade this eval run:
    
    - Expectations: ["assertion 1", "assertion 2", ...]
    - Transcript path: <workspace>/iteration-N/eval-<name>/with_skill/run-1/transcript.md
    - Outputs dir:    <workspace>/iteration-N/eval-<name>/with_skill/run-1/outputs/
    
    Save grading.json to <workspace>/iteration-N/eval-<name>/with_skill/run-1/grading.json.

    Required field names in grading.json: text, passed, evidence (not name/met/details -- the viewer depends on these exact names). For assertions that can be checked programmatically, write and run a script.

  2. Aggregate -- run:

    bash
    python -m scripts.aggregate_benchmark <workspace>/iteration-N --skill-name <name>

    Produces benchmark.json and benchmark.md. List with_skill before baseline. If generating benchmark.json manually, see references/schemas.md for the exact schema the viewer expects.

  3. Analyst pass -- read the benchmark data and surface patterns the aggregate stats might hide: non-discriminating assertions (always pass regardless of skill), high-variance evals (possibly flaky), time/token tradeoffs. See agents/analyzer.md ("Analyzing Benchmark Results" section).

  4. Launch the viewer:

    bash
    nohup python <skill-creator-path>/eval-viewer/generate_review.py \
      <workspace>/iteration-N \
      --skill-name "my-skill" \
      --benchmark <workspace>/iteration-N/benchmark.json \
      > /dev/null 2>&1 &
    VIEWER_PID=$!

    For iteration 2+, add --previous-workspace <workspace>/iteration-<N-1>.

    No display / headless: use --static <output_path> to write a standalone HTML file. The "Submit All Reviews" button downloads feedback.json -- copy it into the workspace directory when done.

  5. Tell the user: "I've opened the results in your browser. 'Outputs' tab lets you review each test case and leave feedback. 'Benchmark' tab shows the quantitative comparison. Come back when you're done."

What the user sees

Outputs tab: one test case at a time -- prompt, output, previous output (collapsed, iteration 2+), formal grades (collapsed), feedback textbox, previous feedback.

Benchmark tab: pass rates, timing, token usage per configuration, per-eval breakdowns, analyst observations.

Navigation: prev/next buttons or arrow keys. "Submit All Reviews" saves feedback.json.

Step 5: Read the feedback
json
{
  "reviews": [
    {"run_id": "eval-0-with_skill", "feedback": "chart is missing axis labels"},
    {"run_id": "eval-1-with_skill", "feedback": ""},
    {"run_id": "eval-2-with_skill", "feedback": "perfect, love this"}
  ],
  "status": "complete"
}

Empty feedback means the user was satisfied. Focus improvements on cases with specific complaints. Kill the viewer when done: kill $VIEWER_PID 2>/dev/null


Improving the skill

This is the heart of the loop. Tests have been run, the user has reviewed -- now make the skill better.

Show full SKILL.md (1,006 more words)Show less
How to think about improvements

Generalize from feedback, do not overfit to examples. The skill will run across countless different prompts. A skill that works only for the test examples is useless. Rather than fiddly constraints, try different metaphors, different working patterns, different framings -- cheap to try, potentially transformative.

Keep the prompt lean. Read the transcripts, not just final outputs. If the skill makes the model waste time on unproductive steps, remove those instructions. Deadweight is not neutral -- it adds noise and slows execution.

Explain the why, not just the what. Models have good theory of mind. They perform better with reasoning than rules. If ALWAYS or NEVER in all caps appears, try instead to explain why the thing matters. "Show the loading state before fetching so users aren't left wondering if anything happened" outperforms "ALWAYS show loading state."

Bundle repeated work. If all test case transcripts show the subagent writing the same helper script, write it once, put it in scripts/, point the skill at it. Every future invocation benefits.

Take time here. Read the draft revision with fresh eyes. Try to genuinely understand what the user wants and what is blocking them, then transmit that understanding into the instructions.

The iteration loop
  1. Apply improvements to the skill
  2. Rerun all test cases into iteration-<N+1>/, including baselines
  3. Launch the viewer with --previous-workspace pointing at the previous iteration
  4. Wait for the user to review
  5. Read feedback, improve again, repeat

Stop when: the user says they are happy, feedback is all empty, or meaningful progress has stopped.


Advanced: Blind comparison

For rigorous comparison between two skill versions, use the blind comparison system: read agents/comparator.md and agents/analyzer.md. An independent agent judges quality without knowing which version is which. Most users do not need this -- the human review loop is usually sufficient.


Description Optimization

The description field is the primary mechanism controlling whether Claude invokes a skill. After creating or improving a skill, offer to optimize the description for triggering accuracy.

Step 1: Generate trigger eval queries

Create 20 queries -- a mix of should-trigger and should-not-trigger. Save as JSON:

json
[
  {"query": "the user prompt", "should_trigger": true},
  {"query": "another prompt", "should_trigger": false}
]

Make queries realistic and specific. Include file paths, personal context, column names, company names, casual speech, typos, abbreviations, varying lengths. Focus on edge cases, not clear-cut cases.

Bad: "Format this data", "Create a chart"

Good: "ok so my boss just sent me this xlsx file (its in my downloads, called something like 'Q4 sales final FINAL v2.xlsx') and she wants me to add a column that shows the profit margin as a percentage. The revenue is in column C and costs are in column D i think"

Should-trigger queries (8-10): Cover different phrasings of the same intent. Include cases where the user clearly needs the skill but does not name it. Include uncommon use cases and cases where this skill competes with another but should win.

Should-not-trigger queries (8-10): The most valuable are near-misses -- queries that share keywords but need something different. Adjacent domains, ambiguous phrasing where keyword-matching would fire but should not. Do not use obviously irrelevant negatives -- they test nothing.

Step 2: Review with user
  1. Read assets/eval_review.html
  2. Replace __EVAL_DATA_PLACEHOLDER__ with the JSON array (no quotes -- JS variable assignment)
  3. Replace __SKILL_NAME_PLACEHOLDER__ and __SKILL_DESCRIPTION_PLACEHOLDER__
  4. Write to /tmp/eval_review_<skill-name>.html and open it
  5. User can edit queries, toggle should-trigger, add/remove entries, then "Export Eval Set"
  6. Check Downloads for the most recent eval_set.json
Step 3: Run the optimization loop
bash
python -m scripts.run_loop \
  --eval-set <path-to-trigger-eval.json> \
  --skill-path <path-to-skill> \
  --model <model-id-powering-this-session> \
  --max-iterations 5 \
  --verbose

Use the model ID from the system prompt so the test matches what the user actually experiences. Run in background and tail output periodically for updates.

The loop: 60/40 train/test split, evaluate current description (3 runs per query for variance), propose improvements from failures, re-evaluate, pick best by test score (not train, to avoid overfitting), repeat up to 5 iterations.

How triggering works: Claude sees skills as (name, description) pairs and decides whether to consult a skill based on that. It only consults skills for tasks it cannot easily handle alone -- eval queries should be substantive enough that a skill would actually help. Simple one-step queries will not trigger skills even with perfect description matching.

Step 4: Apply the result

Take best_description from the JSON output, update the skill's SKILL.md frontmatter. Show before/after to the user and report scores.


Package and Present

If the present_files tool is available:

bash
python -m scripts.package_skill <path/to/skill-folder>

Present the resulting .skill file to the user for installation.


Claude.ai-specific instructions

Same core workflow, different mechanics:

Running test cases: No subagents. Read SKILL.md, follow instructions to complete each test prompt directly, one at a time. Skip baselines. Less rigorous (the skill author is also running it), but useful as a sanity check -- human review compensates.

Reviewing results: If no browser is available, present results directly in conversation. For file outputs (.docx, .xlsx), save to filesystem and tell the user where to find them. Ask for feedback inline.

Benchmarking: Skip quantitative benchmarking -- baseline comparisons are not meaningful without subagents. Focus on qualitative user feedback.

Description optimization: Requires claude -p CLI. Skip if unavailable.

Blind comparison: Requires subagents. Skip.

Updating an existing skill:

  • Preserve the original name (directory name and name frontmatter field unchanged)
  • The installed skill path may be read-only -- copy to /tmp/skill-name/, edit there, package from the copy
  • Direct writes may fail; stage in /tmp/ first

Cowork-specific instructions

  • Subagents work. If severe timeout problems arise, run test prompts in series.
  • No browser/display: use --static <output_path> for the eval viewer, then provide the link for the user to open locally.
  • Generate the eval viewer before evaluating inputs. See Step 4 above.
  • Feedback: "Submit All Reviews" downloads feedback.json. Read from there.
  • Description optimization (run_loop.py) works fine -- uses claude -p.
  • Save description optimization until the skill is fully done and the user agrees.
  • Updating an existing skill: same instructions as Claude.ai section above.

Track these lifecycle steps in TodoWrite if available. In Cowork, specifically add "Create evals JSON and run eval-viewer/generate_review.py so human can review test cases" as a todo item. For a full list of bundled resources, see the Available Resources table near the top of this file.

© curiositech, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 21 other files (scripts, references, assets) in corpus/meta-skills-experiment/self-improved/skill-creator of curiositech/some_claude_skills.

  • SKILL.md
  • EVALUATION.md
  • LICENSE.txt
  • PROVENANCE.md
  • agents/analyzer.md
  • agents/comparator.md
  • agents/grader.md
  • assets/eval_review.html
  • eval-viewer/generate_review.py
  • eval-viewer/viewer.html
  • references/eval-patterns.md
  • references/schemas.md
  • references/troubleshooting.md
  • scripts/__init__.py
  • scripts/aggregate_benchmark.py
  • scripts/generate_report.py
  • … and 6 more

Open the folder on GitHubat commit 6713fc7

Compare with similar skills

Skill Creator next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Skill Creator compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Skill Creator this skillcuriositech/some_claude_skills243—~5.7kAutomated safety check: PassApache-2.0
Skill CreatorAzure/azqr79589 repos~8.2kAutomated safety check: PassApache-2.0
Claude Code Skill Developer Guidediet103/claude-code-infrastructure-showcase10k11 repos~3.5kAutomated safety check: PassMIT
Darwin Skill Optimizeralchaincyf/darwin-skill6.2k1 repos~4.7kAutomated safety check: PassMIT
Claude Code Command Developmentanthropics/claude-plugins-official38k10 repos~4.8kAutomated safety check: PassApache-2.0
Claude Code Plugin Structureanthropics/claude-plugins-official38k10 repos~3.4kAutomated safety check: PassApache-2.0

Similar skills

  • Skill Creator

    Azure/azqr

    Official

    Create new skills, modify and improve existing skills, and measure skill performance.

    795 GitHub starsUsed in 89 repos~8.2k tokens
    Agent WorkflowsAuto-check passed
  • Claude Code Skill Developer Guide

    diet103/claude-code-infrastructure-showcase

    A guide to creating and managing Claude Code skills with auto-activation: skill-rules.json triggers, hooks, enforcement levels, YAML frontmatter and progressive disclosure.

    10k GitHub starsUsed in 11 repos~3.5k tokens
    Agent WorkflowsAuto-check passed
  • Darwin Skill Optimizer

    alchaincyf/darwin-skill

    Scores SKILL.md files on a nine-dimension rubric, then improves them in a keep-or-revert loop with independent judge agents, test prompts, git history and human checkpoints.

    6.2k GitHub starsUsed in 1 repo~4.7k tokens
    Agent WorkflowsAuto-check passed
  • Claude Code Command Development

    anthropics/claude-plugins-official

    Official

    Explains how to write Claude Code slash commands: Markdown files with YAML frontmatter, arguments, file references, bash context and interactive prompts.

    38k GitHub starsUsed in 10 repos~4.8k tokens
    Agent WorkflowsAuto-check passed
  • Claude Code Plugin Structure

    anthropics/claude-plugins-official

    Official

    Explains the directory layout, plugin.json manifest and component organization of a Claude Code plugin, including auto-discovery and portable paths.

    38k GitHub starsUsed in 10 repos~3.4k tokens
    Agent WorkflowsAuto-check passed
  • Skill Release Gate

    rohitg00/ai-engineering-from-scratch

    Evaluates an Agent Skill bundle before release for structure, trigger quality, artifact improvement, script correctness, safety, installed-tree integrity and host portability.

    66k GitHub stars~1k tokensUpdated yesterday
    Agent WorkflowsAuto-check passed

More from curiositech/some_claude_skills

All 109 skills in this repo
  • Crisis Detection Intervention AI

    curiositech/some_claude_skills

    Detect crisis signals in user content using NLP, mental health sentiment analysis, and safe intervention protocols.

    243 GitHub starsUsed in 3 repos~3.8k tokens
    Auto-check passed
  • Form Validation Architect

    curiositech/some_claude_skills

    End-to-end form handling with react-hook-form, Zod schemas, validation patterns, error messaging, field arrays, and multi-step wizards.

    243 GitHub stars~3.8k tokensUpdated 1 mo ago
    Auto-check passed
  • Competitive Cartographer

    curiositech/some_claude_skills

    Strategic analyst that maps competitive landscapes, identifies white space opportunities, and provides positioning recommendations.

    243 GitHub starsUsed in 1 repo~1.3k tokens
    Auto-check passed
  • GitHub Actions Pipeline Builder

    curiositech/some_claude_skills

    Build production CI/CD pipelines with GitHub Actions. An agent skill from curiositech/some_claude_skills.

    243 GitHub stars~2.8k tokensUpdated 1 mo ago
    Auto-check: notes
  • Computer Vision Pipeline

    curiositech/some_claude_skills

    Build production computer vision pipelines for object detection, tracking, and video analysis.

    243 GitHub starsUsed in 1 repo~4k tokens
    Auto-check passed
  • Design Archivist

    curiositech/some_claude_skills

    Long-running design anthropologist that builds comprehensive visual databases from 500-1000 real-world examples, extracting color palettes, typography patterns, layout systems, and interaction…

    243 GitHub starsUsed in 1 repo~1.6k tokens
    Auto-check passed

Categories

Questions about Skill Creator

What does Skill Creator do?

Guide for creating, improving, benchmarking, and packaging Claude Agent Skills (SKILL.md files). Skill Creator is an agent skill from curiositech/some_claude_skills.md files).

When should I use Skill Creator?

Skill Creator fits situations like: S want to create a skill from scratch; test an existing skill; benchmark skill performance with variance analysis; optimize a skill description for triggering accuracy.

How do I install Skill Creator in Claude Code?

Run `npx skills add curiositech/some_claude_skills --skill skill-creator -a claude-code`. Or copy the skill folder (corpus/meta-skills-experiment/self-improved/skill-creator in curiositech/some_claude_skills) into .claude/skills/skill-creator in your project. Claude Code loads it when a task matches its description.

How do I install Skill Creator in Codex?

Run `npx skills add curiositech/some_claude_skills --skill skill-creator -a codex`. Or copy the skill folder (corpus/meta-skills-experiment/self-improved/skill-creator in curiositech/some_claude_skills) into .agents/skills/skill-creator in your project. Codex loads it when a task matches its description.

Can I use Skill Creator in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add curiositech/some_claude_skills --skill skill-creator -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/skill-creator, .gemini/skills/skill-creator, .github/skills/skill-creator and .opencode/skills/skill-creator in your project.

What does Skill Creator need to run?

Going by SKILL.md and its folder, Skill Creator needs Python for the scripts in its folder and the command-line tools its instructions call (python and claude). Our summary lists: Python 3.

Does Skill Creator access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Skill Creator safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Skill Creator use?

Skill Creator is published under the Apache-2.0 licence (from the LICENSE file in the skill folder). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Skill Creator use?

About 5.7k tokens (SKILL.md is roughly 23k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 6.6k tokens, read only when the agent opens those files.

What are the alternatives to Skill Creator?

Skills that share tags, products or a category with Skill Creator: Skill Creator (Azure/azqr, 795 stars), Claude Code Skill Developer Guide (diet103/claude-code-infrastructure-showcase, 10k stars), Darwin Skill Optimizer (alchaincyf/darwin-skill, 6.2k stars) and Claude Code Command Development (anthropics/claude-plugins-official, 38k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Skill Creator?

curiositech (a GitHub organization) maintains it in curiositech/some_claude_skills, which has 243 GitHub stars. The repository holds 109 skills in this directory. The repository was last updated on September 6, 2026.

Source: curiositech/some_claude_skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.