Skill Creator
Azure/azqr
Create new skills, modify and improve existing skills, and measure skill performance.
Guide for creating, improving, benchmarking, and packaging Claude Agent Skills (SKILL.md files).
$ npx skills add curiositech/some_claude_skills --skill skill-creator -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install curiositech/some_claude_skills skill-creator --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/curiositech/some_claude_skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/corpus/meta-skills-experiment/self-improved/skill-creator .claude/skills/skill-creator && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "skill-creator" agent skill from https://github.com/curiositech/some_claude_skills/tree/main/corpus/meta-skills-experiment/self-improved/skill-creator into .claude/skills/skill-creator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "skill-creator", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/curiositech/some_claude_skills/tree/main/corpus/meta-skills-experiment/self-improved/skill-creatorType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add curiositech/some_claude_skills --skill skill-creator -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install curiositech/some_claude_skills skill-creator --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/curiositech/some_claude_skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/corpus/meta-skills-experiment/self-improved/skill-creator .agents/skills/skill-creator && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "skill-creator" agent skill from https://github.com/curiositech/some_claude_skills/tree/main/corpus/meta-skills-experiment/self-improved/skill-creator into .agents/skills/skill-creator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "skill-creator", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add curiositech/some_claude_skills --skill skill-creator -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install curiositech/some_claude_skills skill-creator --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/curiositech/some_claude_skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/corpus/meta-skills-experiment/self-improved/skill-creator .cursor/skills/skill-creator && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "skill-creator" agent skill from https://github.com/curiositech/some_claude_skills/tree/main/corpus/meta-skills-experiment/self-improved/skill-creator into .cursor/skills/skill-creator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "skill-creator", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/curiositech/some_claude_skills.git --path corpus/meta-skills-experiment/self-improved/skill-creator--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add curiositech/some_claude_skills --skill skill-creator -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install curiositech/some_claude_skills skill-creator --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/curiositech/some_claude_skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/corpus/meta-skills-experiment/self-improved/skill-creator .gemini/skills/skill-creator && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "skill-creator" agent skill from https://github.com/curiositech/some_claude_skills/tree/main/corpus/meta-skills-experiment/self-improved/skill-creator into .gemini/skills/skill-creator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "skill-creator", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install curiositech/some_claude_skills skill-creatorInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add curiositech/some_claude_skills --skill skill-creator -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/curiositech/some_claude_skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/corpus/meta-skills-experiment/self-improved/skill-creator .github/skills/skill-creator && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "skill-creator" agent skill from https://github.com/curiositech/some_claude_skills/tree/main/corpus/meta-skills-experiment/self-improved/skill-creator into .github/skills/skill-creator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "skill-creator", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add curiositech/some_claude_skills --skill skill-creator -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install curiositech/some_claude_skills skill-creator --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/curiositech/some_claude_skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/corpus/meta-skills-experiment/self-improved/skill-creator .opencode/skills/skill-creator && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "skill-creator" agent skill from https://github.com/curiositech/some_claude_skills/tree/main/corpus/meta-skills-experiment/self-improved/skill-creator into .opencode/skills/skill-creator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "skill-creator", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
skill-creatorGuide for creating, improving, benchmarking, and packaging Claude Agent Skills (SKILL.md files).
Skill Creator is an agent skill from curiositech/some_claude_skills. Guide for creating, improving, benchmarking, and packaging Claude Agent Skills (SKILL.md files). Invoke when users want to create a skill from scratch, improve or test an existing skill, benchmark skill performance with variance analysis, or optimize a skill description for triggering accuracy. Also invoke when users say "turn this into a skill", "make a skill for X", "help me write a SKILL.md", "my skill isn't firing correctly", or want to convert a workflow/conversation into a reusable skill. Invoke proactively…
Its SKILL.md is about 5.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 26 other files, including scripts, reference files and assets (for example `EVALUATION.md`, `PROVENANCE.md` and `agents/analyzer.md`).
It sits in Agent Workflows, covering Skill authoring. The repository describes itself as: Claude skills that make my life easier. The licence is Apache-2.0.
9 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 6713fc7. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 3 files in scripts/ (Python, from the files we listed), which the agent can run.
Shell commands in SKILL.md call:
pythonclaudeFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Skill Creator loads about 5.7k tokens when it runs, and up to ~12k if it reads all its reference files. Until then it costs about 177 tokens; SKILL.md has 2,481 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from curiositech/some_claude_skills at commit 6713fc7, republished under its Apache-2.0 licence (© curiositech). 2,481 words, ~5,747 tokens.
.claude/skills/skill-creator/SKILL.md (or your agent's skills folder). This skill also uses 21 other files; get the full folder from GitHub.Create new skills and iteratively improve them. The lifecycle:
Determine where the user is in this lifecycle and help them progress. Starting from scratch ("I want a skill for X") -- help narrow scope, draft, test, iterate. Already have a draft -- go straight to eval/iterate. Just want a better description -- jump to Description Optimization.
Be flexible: if the user says "just vibe with me, no evals," do that.
Before starting, verify what is available in the environment. The full workflow requires these bundled resources:
| Resource | Path | Used for |
|---|---|---|
| Grader instructions | agents/grader.md | Assertion evaluation |
| Comparator instructions | agents/comparator.md | Blind A/B comparison |
| Analyzer instructions | agents/analyzer.md | Benchmark pattern analysis |
| Schema reference | references/schemas.md | evals.json, grading.json formats |
| Assertion patterns | references/eval-patterns.md | Writing discriminating assertions |
| Troubleshooting | references/troubleshooting.md | Common errors and fixes |
| Eval review template | assets/eval_review.html | Description optimization UI |
| Benchmark aggregator | scripts/aggregate_benchmark | python -m scripts.aggregate_benchmark |
| Description optimizer | scripts/run_loop | python -m scripts.run_loop |
| Eval viewer | eval-viewer/generate_review.py | Human review interface |
If resources are missing: The core loop (draft -> test -> review -> improve) still works without them -- grade inline and present results in-conversation instead of the browser viewer. Note what is unavailable to the user and adapt.
The skill creator serves people across a wide range of technical backgrounds. Non-developers are opening terminals for the first time because Claude makes things possible; experienced engineers are also common.
Pay attention to context cues:
Err toward clarity. A short inline definition never hurts.
The current conversation may already contain the workflow to capture. If the user says "turn this into a skill" or "make this repeatable," mine the conversation history first:
Fill gaps with the user and confirm before drafting.
If starting from scratch, establish:
Ask about edge cases, input/output examples, success criteria, and dependencies before writing test prompts. If useful MCPs are available (docs search, similar skill lookup), research in parallel via subagents. Come prepared so the interview is lightweight.
Based on the interview, compose the skill. Every skill has:
Use this as the starting structure -- fill in or remove sections as needed:
---
name: skill-name
description: >
[What this skill does and when to trigger it. One to three sentences.
Apply the pushy principle: include adjacent trigger contexts.]
---
# [Skill Name]
[One paragraph: what this skill enables and when to use it.]
## [Core workflow or main section]
[Instructions in imperative form. Explain the *why* behind each step.]
## Output format
[Concrete template for what this skill produces.]
## Examples
**Example 1:**
Input: [realistic user prompt]
Output: [expected result]
## Reference files
- Read `references/advanced.md` when [specific situation]
- Run `scripts/transform.py` for [repeatable operation]skill-name/
+-- SKILL.md (required)
+-- Optional bundled resources:
+-- scripts/ - Scripts for deterministic/repetitive tasks
+-- references/ - Docs loaded into context as needed
+-- assets/ - Templates, icons, fontsProgressive loading:
Keep SKILL.md under 500 lines. If approaching that limit, add a layer of hierarchy with clear pointers to reference files. For reference files over 300 lines, add a table of contents.
Domain organization -- when a skill covers multiple frameworks/variants:
cloud-deploy/
+-- SKILL.md (workflow + selection logic)
+-- references/
+-- aws.md
+-- gcp.md
+-- azure.mdClaude reads only the relevant reference file.
Output templates -- be concrete:
## Report structure
ALWAYS use this template:
# [Title]
## Executive summary
## Key findings
## RecommendationsExamples -- include realistic ones:
## Commit message format
**Example:**
Input: Added user authentication with JWT tokens
Output: feat(auth): implement JWT-based authenticationWriting style -- use imperative form. Explain the why rather than issuing rules. Avoid ALWAYS/NEVER in all caps where the reasoning can be explained instead -- models respond better to understanding than mandates. Write a draft, then read it fresh and improve it.
Security -- skills must not contain malware, exploit code, or anything that would surprise the user if they read the description. Do not create skills designed to facilitate unauthorized access or data exfiltration. Roleplay skills are fine.
After drafting, write 2-3 realistic test prompts -- things a real user would actually type. Share them: "Here are a few test cases I'd like to try. Do these look right, or do you want to add more?"
Save to evals/evals.json. Do not write assertions yet -- just prompts.
Assertions get drafted while runs are in progress.
{
"skill_name": "example-skill",
"evals": [
{
"id": 1,
"prompt": "User's task prompt",
"expected_output": "Description of expected result",
"files": []
}
]
}See references/schemas.md for the full schema including the assertions field.
This is one continuous sequence -- do not stop partway through. Do NOT use
/skill-test or any other testing skill.
Put results in <skill-name>-workspace/ as a sibling to the skill directory.
Organize by iteration, then by eval name, then by configuration, then by run.
Create directories as you go, not all upfront.
<skill-name>-workspace/
+-- iteration-1/
| +-- eval-<name>/ <-- one per test case (descriptive name)
| | +-- eval_metadata.json <-- eval ID, prompt, assertions
| | +-- with_skill/
| | | +-- run-1/
| | | +-- outputs/ <-- files the subagent produced
| | | | +-- metrics.json
| | | +-- timing.json <-- from task notification
| | | +-- grading.json <-- written by grader agent
| | +-- without_skill/ <-- or old_skill/ when improving
| | +-- run-1/
| | +-- outputs/
| | +-- timing.json
| | +-- grading.json
| +-- benchmark.json <-- written by aggregate_benchmark.py
| +-- benchmark.md
+-- iteration-2/
+-- ...For each test case, spawn two subagents simultaneously -- one with the skill, one without. Do not do with-skill runs first and baseline runs later. Launch everything at once so results arrive around the same time.
With-skill run:
Execute this task:
- Skill path: <path-to-skill>
- Task: <eval prompt>
- Input files: <eval files if any, or "none">
- Save outputs to: <workspace>/iteration-<N>/eval-<name>/with_skill/run-1/outputs/
- Outputs to save: <what the user cares about>Baseline run -- depends on context:
<workspace>/iteration-<N>/eval-<name>/without_skill/run-1/outputs/cp -r <skill-path> <workspace>/skill-snapshot/), point baseline at snapshot,
save to old_skill/run-1/outputs/Write eval_metadata.json at the eval level (e.g.,
<workspace>/iteration-N/eval-<name>/eval_metadata.json). Give each eval a
descriptive name based on what it tests -- not just "eval-0":
{
"eval_id": 0,
"eval_name": "descriptive-name-here",
"prompt": "The user's task prompt",
"assertions": []
}Do not just wait for runs to finish -- use this time productively. Draft
quantitative assertions for each test case and explain them to the user. If
assertions already exist in evals/evals.json, review and explain them.
Good assertions are objectively verifiable and have descriptive names -- someone
glancing at benchmark results should immediately understand what each one checks.
For subjective skills (writing style, design quality), skip assertions and rely
on qualitative review. See references/eval-patterns.md for assertion patterns
by task type and the discriminating assertion test.
Update eval_metadata.json files and evals/evals.json with assertions once
drafted. Preview for the user: explain what they will see in the viewer.
When each subagent task completes, save timing data immediately to timing.json
in the run directory (e.g., with_skill/run-1/timing.json):
{
"total_tokens": 84852,
"duration_ms": 23332,
"total_duration_seconds": 23.3
}This data only comes through task completion notifications -- capture it then.
GENERATE THE EVAL VIEWER BEFORE EVALUATING INPUTS YOURSELF. Get outputs in front of the human first.
Once all runs complete:
Grade each run -- spawn a grader subagent for each run directory (both with_skill and without_skill/old_skill get their own grader). Use this prompt:
You are a grader agent. Read agents/grader.md from <skill-creator-path>
to load your full instructions, then grade this eval run:
- Expectations: ["assertion 1", "assertion 2", ...]
- Transcript path: <workspace>/iteration-N/eval-<name>/with_skill/run-1/transcript.md
- Outputs dir: <workspace>/iteration-N/eval-<name>/with_skill/run-1/outputs/
Save grading.json to <workspace>/iteration-N/eval-<name>/with_skill/run-1/grading.json.Required field names in grading.json: text, passed, evidence (not
name/met/details -- the viewer depends on these exact names). For
assertions that can be checked programmatically, write and run a script.
Aggregate -- run:
python -m scripts.aggregate_benchmark <workspace>/iteration-N --skill-name <name>Produces benchmark.json and benchmark.md. List with_skill before baseline.
If generating benchmark.json manually, see references/schemas.md for the
exact schema the viewer expects.
Analyst pass -- read the benchmark data and surface patterns the aggregate
stats might hide: non-discriminating assertions (always pass regardless of
skill), high-variance evals (possibly flaky), time/token tradeoffs. See
agents/analyzer.md ("Analyzing Benchmark Results" section).
Launch the viewer:
nohup python <skill-creator-path>/eval-viewer/generate_review.py \
<workspace>/iteration-N \
--skill-name "my-skill" \
--benchmark <workspace>/iteration-N/benchmark.json \
> /dev/null 2>&1 &
VIEWER_PID=$!For iteration 2+, add --previous-workspace <workspace>/iteration-<N-1>.
No display / headless: use --static <output_path> to write a standalone
HTML file. The "Submit All Reviews" button downloads feedback.json -- copy
it into the workspace directory when done.
Tell the user: "I've opened the results in your browser. 'Outputs' tab lets you review each test case and leave feedback. 'Benchmark' tab shows the quantitative comparison. Come back when you're done."
Outputs tab: one test case at a time -- prompt, output, previous output (collapsed, iteration 2+), formal grades (collapsed), feedback textbox, previous feedback.
Benchmark tab: pass rates, timing, token usage per configuration, per-eval breakdowns, analyst observations.
Navigation: prev/next buttons or arrow keys. "Submit All Reviews" saves
feedback.json.
{
"reviews": [
{"run_id": "eval-0-with_skill", "feedback": "chart is missing axis labels"},
{"run_id": "eval-1-with_skill", "feedback": ""},
{"run_id": "eval-2-with_skill", "feedback": "perfect, love this"}
],
"status": "complete"
}Empty feedback means the user was satisfied. Focus improvements on cases with
specific complaints. Kill the viewer when done: kill $VIEWER_PID 2>/dev/null
This is the heart of the loop. Tests have been run, the user has reviewed -- now make the skill better.
Generalize from feedback, do not overfit to examples. The skill will run across countless different prompts. A skill that works only for the test examples is useless. Rather than fiddly constraints, try different metaphors, different working patterns, different framings -- cheap to try, potentially transformative.
Keep the prompt lean. Read the transcripts, not just final outputs. If the skill makes the model waste time on unproductive steps, remove those instructions. Deadweight is not neutral -- it adds noise and slows execution.
Explain the why, not just the what. Models have good theory of mind. They perform better with reasoning than rules. If ALWAYS or NEVER in all caps appears, try instead to explain why the thing matters. "Show the loading state before fetching so users aren't left wondering if anything happened" outperforms "ALWAYS show loading state."
Bundle repeated work. If all test case transcripts show the subagent writing
the same helper script, write it once, put it in scripts/, point the skill at
it. Every future invocation benefits.
Take time here. Read the draft revision with fresh eyes. Try to genuinely understand what the user wants and what is blocking them, then transmit that understanding into the instructions.
iteration-<N+1>/, including baselines--previous-workspace pointing at the previous
iterationStop when: the user says they are happy, feedback is all empty, or meaningful progress has stopped.
For rigorous comparison between two skill versions, use the blind comparison
system: read agents/comparator.md and agents/analyzer.md. An independent
agent judges quality without knowing which version is which. Most users do not
need this -- the human review loop is usually sufficient.
The description field is the primary mechanism controlling whether Claude
invokes a skill. After creating or improving a skill, offer to optimize the
description for triggering accuracy.
Create 20 queries -- a mix of should-trigger and should-not-trigger. Save as JSON:
[
{"query": "the user prompt", "should_trigger": true},
{"query": "another prompt", "should_trigger": false}
]Make queries realistic and specific. Include file paths, personal context, column names, company names, casual speech, typos, abbreviations, varying lengths. Focus on edge cases, not clear-cut cases.
Bad: "Format this data", "Create a chart"
Good: "ok so my boss just sent me this xlsx file (its in my downloads, called something like 'Q4 sales final FINAL v2.xlsx') and she wants me to add a column that shows the profit margin as a percentage. The revenue is in column C and costs are in column D i think"
Should-trigger queries (8-10): Cover different phrasings of the same intent. Include cases where the user clearly needs the skill but does not name it. Include uncommon use cases and cases where this skill competes with another but should win.
Should-not-trigger queries (8-10): The most valuable are near-misses -- queries that share keywords but need something different. Adjacent domains, ambiguous phrasing where keyword-matching would fire but should not. Do not use obviously irrelevant negatives -- they test nothing.
assets/eval_review.html__EVAL_DATA_PLACEHOLDER__ with the JSON array (no quotes -- JS
variable assignment)__SKILL_NAME_PLACEHOLDER__ and __SKILL_DESCRIPTION_PLACEHOLDER__/tmp/eval_review_<skill-name>.html and open iteval_set.jsonpython -m scripts.run_loop \
--eval-set <path-to-trigger-eval.json> \
--skill-path <path-to-skill> \
--model <model-id-powering-this-session> \
--max-iterations 5 \
--verboseUse the model ID from the system prompt so the test matches what the user actually experiences. Run in background and tail output periodically for updates.
The loop: 60/40 train/test split, evaluate current description (3 runs per query for variance), propose improvements from failures, re-evaluate, pick best by test score (not train, to avoid overfitting), repeat up to 5 iterations.
How triggering works: Claude sees skills as (name, description) pairs and
decides whether to consult a skill based on that. It only consults skills for
tasks it cannot easily handle alone -- eval queries should be substantive enough
that a skill would actually help. Simple one-step queries will not trigger skills
even with perfect description matching.
Take best_description from the JSON output, update the skill's SKILL.md
frontmatter. Show before/after to the user and report scores.
If the present_files tool is available:
python -m scripts.package_skill <path/to/skill-folder>Present the resulting .skill file to the user for installation.
Same core workflow, different mechanics:
Running test cases: No subagents. Read SKILL.md, follow instructions to complete each test prompt directly, one at a time. Skip baselines. Less rigorous (the skill author is also running it), but useful as a sanity check -- human review compensates.
Reviewing results: If no browser is available, present results directly in conversation. For file outputs (.docx, .xlsx), save to filesystem and tell the user where to find them. Ask for feedback inline.
Benchmarking: Skip quantitative benchmarking -- baseline comparisons are not meaningful without subagents. Focus on qualitative user feedback.
Description optimization: Requires claude -p CLI. Skip if unavailable.
Blind comparison: Requires subagents. Skip.
Updating an existing skill:
name frontmatter field
unchanged)/tmp/skill-name/, edit
there, package from the copy/tmp/ first--static <output_path> for the eval viewer, then
provide the link for the user to open locally.feedback.json. Read from there.run_loop.py) works fine -- uses claude -p.Track these lifecycle steps in TodoWrite if available. In Cowork, specifically
add "Create evals JSON and run eval-viewer/generate_review.py so human can
review test cases" as a todo item. For a full list of bundled resources, see the
Available Resources table near the top of this file.
© curiositech, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 21 other files (scripts, references, assets) in corpus/meta-skills-experiment/self-improved/skill-creator of curiositech/some_claude_skills.
Open the folder on GitHubat commit 6713fc7
Skill Creator next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Skill Creator this skillcuriositech/some_claude_skills | 243 | — | ~5.7k | Automated safety check: Pass | Apache-2.0 | |
| Skill CreatorAzure/azqr | 795 | 89 repos | ~8.2k | Automated safety check: Pass | Apache-2.0 | |
| Claude Code Skill Developer Guidediet103/claude-code-infrastructure-showcase | 10k | 11 repos | ~3.5k | Automated safety check: Pass | MIT | |
| Darwin Skill Optimizeralchaincyf/darwin-skill | 6.2k | 1 repos | ~4.7k | Automated safety check: Pass | MIT | |
| Claude Code Command Developmentanthropics/claude-plugins-official | 38k | 10 repos | ~4.8k | Automated safety check: Pass | Apache-2.0 | |
| Claude Code Plugin Structureanthropics/claude-plugins-official | 38k | 10 repos | ~3.4k | Automated safety check: Pass | Apache-2.0 |
Azure/azqr
Create new skills, modify and improve existing skills, and measure skill performance.
diet103/claude-code-infrastructure-showcase
A guide to creating and managing Claude Code skills with auto-activation: skill-rules.json triggers, hooks, enforcement levels, YAML frontmatter and progressive disclosure.
alchaincyf/darwin-skill
Scores SKILL.md files on a nine-dimension rubric, then improves them in a keep-or-revert loop with independent judge agents, test prompts, git history and human checkpoints.
anthropics/claude-plugins-official
Explains how to write Claude Code slash commands: Markdown files with YAML frontmatter, arguments, file references, bash context and interactive prompts.
anthropics/claude-plugins-official
Explains the directory layout, plugin.json manifest and component organization of a Claude Code plugin, including auto-discovery and portable paths.
rohitg00/ai-engineering-from-scratch
Evaluates an Agent Skill bundle before release for structure, trigger quality, artifact improvement, script correctness, safety, installed-tree integrity and host portability.
curiositech/some_claude_skills
Detect crisis signals in user content using NLP, mental health sentiment analysis, and safe intervention protocols.
curiositech/some_claude_skills
End-to-end form handling with react-hook-form, Zod schemas, validation patterns, error messaging, field arrays, and multi-step wizards.
curiositech/some_claude_skills
Strategic analyst that maps competitive landscapes, identifies white space opportunities, and provides positioning recommendations.
curiositech/some_claude_skills
Build production CI/CD pipelines with GitHub Actions. An agent skill from curiositech/some_claude_skills.
curiositech/some_claude_skills
Build production computer vision pipelines for object detection, tracking, and video analysis.
curiositech/some_claude_skills
Long-running design anthropologist that builds comprehensive visual databases from 500-1000 real-world examples, extracting color palettes, typography patterns, layout systems, and interaction…
Categories
Guide for creating, improving, benchmarking, and packaging Claude Agent Skills (SKILL.md files). Skill Creator is an agent skill from curiositech/some_claude_skills.md files).
Skill Creator fits situations like: S want to create a skill from scratch; test an existing skill; benchmark skill performance with variance analysis; optimize a skill description for triggering accuracy.
Run `npx skills add curiositech/some_claude_skills --skill skill-creator -a claude-code`. Or copy the skill folder (corpus/meta-skills-experiment/self-improved/skill-creator in curiositech/some_claude_skills) into .claude/skills/skill-creator in your project. Claude Code loads it when a task matches its description.
Run `npx skills add curiositech/some_claude_skills --skill skill-creator -a codex`. Or copy the skill folder (corpus/meta-skills-experiment/self-improved/skill-creator in curiositech/some_claude_skills) into .agents/skills/skill-creator in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add curiositech/some_claude_skills --skill skill-creator -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/skill-creator, .gemini/skills/skill-creator, .github/skills/skill-creator and .opencode/skills/skill-creator in your project.
Going by SKILL.md and its folder, Skill Creator needs Python for the scripts in its folder and the command-line tools its instructions call (python and claude). Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Skill Creator is published under the Apache-2.0 licence (from the LICENSE file in the skill folder). It allows redistribution, so the full SKILL.md is shown on this page.
About 5.7k tokens (SKILL.md is roughly 23k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 6.6k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Skill Creator: Skill Creator (Azure/azqr, 795 stars), Claude Code Skill Developer Guide (diet103/claude-code-infrastructure-showcase, 10k stars), Darwin Skill Optimizer (alchaincyf/darwin-skill, 6.2k stars) and Claude Code Command Development (anthropics/claude-plugins-official, 38k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
curiositech (a GitHub organization) maintains it in curiositech/some_claude_skills, which has 243 GitHub stars. The repository holds 109 skills in this directory. The repository was last updated on September 6, 2026.
Source: curiositech/some_claude_skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.