Agent skill

Hk Skill Creator

by deepklarity in deepklarity/harness-kit

Create new skills, modify and improve existing skills, and measure skill performance.

MITAuto-check: notesAgent Workflows

Install Hk Skill Creator

skills CLI
$ npx skills add deepklarity/harness-kit --skill hk-skill-creator -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install deepklarity/harness-kit hk-skill-creator --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/deepklarity/harness-kit.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/hk-skill-creator .claude/skills/hk-skill-creator && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
hk-skill-creator
GitHub stars
100
Token cost
~3.2k tokens
SKILL.md length
1,587 words
Files
4 (incl. scripts, references)
Skills in repo
18
Repo updated
First seen
Licence
MIT

At a glance

Create new skills, modify and improve existing skills, and measure skill performance.

  • Works in 9 steps: Capture Intent → Plan Resources → Initialize or Locate the Skill → …
  • Users want to create a skill from scratch
  • SKILL.md covers Context, The Core Loop, Communication Style and Step 1: Capture Intent, plus 9 more sections
  • Runs Python scripts from its folder; calls python

What it does

Hk Skill Creator is an agent skill from deepklarity/harness-kit. Create new skills, modify and improve existing skills, and measure skill performance. Use when users want to create a skill from scratch, edit or optimize an existing skill, run evals to test a skill, benchmark skill performance with variance analysis, or optimize a skill's description for better triggering accuracy. Triggers on: 'create a skill', 'new skill', 'make a skill for', 'improve this skill', 'test this skill', 'skill eval', 'optimize skill description', or /hk-skill-creator.

Its SKILL.md is about 3.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including scripts and reference files (for example `references/writing_guide.md`, `scripts/init_skill.py` and `scripts/validate_skill.py`).

It sits in Agent Workflows, covering Skill authoring and LLM evaluation. The repository describes itself as: A kit for building with AI agents and also the engineering patterns around it. The licence is MIT.

When your agent uses it

  • Users want to create a skill from scratch
  • Optimize an existing skill
  • Run evals to test a skill
  • Benchmark skill performance with variance analysis

Example prompts

  • “s description for better triggering accuracy. Triggers on:”
  • “new skill”
  • “make a skill for”
  • “/hk-skill-creator”

Requirements

  • Python 3
  • Pre-approved tools (allowed-tools): Bash, Read, Edit, Write, Grep, Glob, Agent, Task, TaskCreate, TaskUpdate

Workflow steps

9 steps, taken from the step headings in SKILL.md.

  1. Capture Intent
  2. Plan Resources
  3. Initialize or Locate the Skill
  4. Write the Skill
  5. Test the Skill
  6. Evaluate Results
  7. Improve the Skill
  8. Validate
  9. Optimize Description (Optional)

What it can do on your machine

Read from SKILL.md and the folder at commit 87305cd. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Bash
    • Read
    • Edit
    • Write
    • Grep
    • Glob
    • Agent
    • Task
    • TaskCreate
    • TaskUpdate

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Hk Skill Creator loads about 3.2k tokens when it runs, and up to ~6.5k if it reads all its reference files. Until then it costs about 127 tokens; SKILL.md has 1,587 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~127
When it runs · the whole SKILL.md, loaded when a task matches
~3.2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~6.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Bash, Read, Edit, Write, Grep, Glob, Agent, Task, TaskCreate, TaskUpdate

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from deepklarity/harness-kit at commit 87305cd, republished under its MIT licence (© deepklarity). 1,587 words, ~3,177 tokens.

Download SKILL.mdSave it as .claude/skills/hk-skill-creator/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
hk-skill-creator
description
Create new skills, modify and improve existing skills, and measure skill performance. Use when users want to create a skill from scratch, edit or optimize an existing skill, run evals to test a skill, benchmark skill performance with variance analysis, or optimize a skill's description for better triggering accuracy. Triggers on: 'create a skill', 'new skill', 'make a skill for', 'improve this skill', 'test this skill', 'skill eval', 'optimize skill description', or /hk-skill-creator.
allowed-tools
Bash, Read, Edit, Write, Grep, Glob, Agent, Task, TaskCreate, TaskUpdate
argument-hint
[optional: what the skill should do, or path to existing skill to improve]

/hk-skill-creator — Create and Improve Skills

Create new skills and iteratively improve existing ones through a structured draft-test-evaluate-improve loop.

Context

<skill_context> $ARGUMENTS </skill_context>

If the context above describes what skill to create or improve, proceed. If empty or unclear, ask the user what they want the skill to do.

The Core Loop

The process of creating a skill:

  1. Understand what the skill should do, with concrete examples
  2. Draft the SKILL.md and any bundled resources
  3. Test by running Claude-with-the-skill on realistic prompts
  4. Evaluate outputs qualitatively (does it do what the user wants?) and quantitatively (assertions)
  5. Improve based on feedback — generalize, don't overfit
  6. Repeat until the user is satisfied
  7. Optimize the description for reliable triggering

Figure out where the user is in this process and help them progress. Maybe they want a skill from scratch — help narrow intent, draft, test. Maybe they already have a draft — go straight to eval/iterate. Maybe they just want to vibe — be flexible.

Communication Style

Skills are used by people across a wide range of technical familiarity. Pay attention to context cues:

  • "evaluation" and "benchmark" are borderline but OK to use without explanation
  • For "JSON", "assertion", "frontmatter" — look for cues the user knows these before using them without a brief definition
  • When in doubt, briefly explain terms inline
  • Match the user's register — if they're casual, be casual back

Step 1: Capture Intent

Start by understanding what the skill should enable. The conversation might already contain a workflow to capture (e.g., "turn this into a skill"). If so, extract answers from context first — tools used, sequence of steps, corrections made, input/output formats. The user fills gaps, then confirms before proceeding.

Questions to resolve:

  • What should this skill enable Claude to do?
  • When should this skill trigger? (what user phrases/contexts)
  • What's the expected output format or behavior?
  • Are there concrete examples of input/output?
Interview and Research

Proactively ask about edge cases, input/output formats, example files, success criteria, and dependencies. Check existing skills in .claude/skills/ — spawn a haiku subagent to scan for overlapping or related skills. Come prepared with context to reduce burden on the user.

Do not proceed to drafting until the functionality is clearly understood.

Step 2: Plan Resources

For each concrete example from Step 1, analyze:

  1. How to execute from scratch (what steps, what tools)
  2. What scripts, references, or assets would help when executing repeatedly

Build a list of reusable resources to include. Common signals:

  • If all test cases would independently write similar helper scripts, bundle that script in scripts/
  • If the skill needs detailed reference docs (API schemas, large examples), put them in references/
  • If the skill produces output from templates, put templates in assets/

Step 3: Initialize or Locate the Skill

New skill

Scaffold with:

bash
python ${CLAUDE_SKILL_DIR}/scripts/init_skill.py <skill-name>

This creates the directory structure at .claude/skills/<skill-name>/ with a template SKILL.md. Follow hk- prefix convention for this project.

Existing skill

Read the existing SKILL.md and understand its current structure before making changes.

Step 4: Write the Skill

The skill is written for another instance of Claude to use. Focus on non-obvious procedural knowledge — things the model wouldn't know or would get wrong without guidance.

Implementation order
  1. Bundled resources first (scripts/, references/, assets/)
  2. Delete unused example files from initialization
  3. Complete SKILL.md last (it references the resources)
SKILL.md Structure

Read references/writing_guide.md for the complete writing guide including all frontmatter fields, dynamic context injection, and subagent execution patterns. Key points:

Frontmatter — The description is the primary triggering mechanism. Descriptions longer than 250 characters get truncated in the skill listing, so front-load the key use case. Include both what the skill does AND specific trigger contexts. Descriptions should be slightly "pushy" to combat undertriggering.

Invocation control — Two fields control who can trigger a skill:

  • disable-model-invocation: true — only the user can invoke via /name (use for side-effect workflows like deploy, commit)
  • user-invocable: false — only Claude can invoke (use for background knowledge, not actionable as a command)

Subagent execution — Set context: fork to run the skill in an isolated subagent. Pair with agent: to pick the agent type (Explore, Plan, general-purpose, or a custom agent from .claude/agents/). Only use for task-oriented skills with explicit instructions — guidelines-only skills produce nothing useful in a fork because the subagent has no conversation context.

Dynamic context injection — Use !`command` syntax to run shell commands before the skill content reaches Claude. The output replaces the placeholder. Useful for injecting live data (git status, API state, file listings).

String substitutions — $ARGUMENTS for all args, $ARGUMENTS[N] or $N for positional access, ${CLAUDE_SESSION_ID} for session ID, ${CLAUDE_SKILL_DIR} for the skill's directory path.

Body — Use imperative form. Explain why things matter, not just what to do. Today's LLMs are smart — when given reasoning they can generalize beyond rote instructions. If you find yourself writing ALWAYS or NEVER in all caps, reframe with reasoning instead. Keep under 500 lines; move detailed content to references/.

Conventions for this project:

  • Follow patterns in existing .claude/skills/hk-* skills
  • Include $ARGUMENTS context block if the skill accepts arguments
  • Use argument-hint: in frontmatter to show expected arguments
  • Use ${CLAUDE_SKILL_DIR} to reference bundled scripts/files portably
  • Reference all bundled resources explicitly with guidance on when to read them

Step 5: Test the Skill

Create test prompts

Come up with 2-3 realistic test prompts — the kind of thing a real user would actually say. Not abstract requests, but concrete and specific with detail (file paths, personal context, column names, backstory). Share with the user for approval before running.

Save test cases to a workspace directory:

<skill-name>-workspace/
├── evals.json              # Test prompts and expected behaviors
├── iteration-1/
│   ├── eval-<name>/
│   │   ├── with_skill/     # Output from run with skill
│   │   └── without_skill/  # Baseline output (no skill)
│   └── ...
└── iteration-2/
    └── ...
Run test cases

For each test case, spawn two subagents in the same message — one with the skill loaded, one without (baseline). Launch all runs at once for parallel execution.

With-skill run (sonnet subagent):

Execute this task following the skill at <path-to-skill>/SKILL.md:
- Task: <eval prompt>
- Input files: <if any>
- Save outputs to: <workspace>/iteration-N/eval-<name>/with_skill/

Baseline run (sonnet subagent, same prompt, no skill):

Execute this task WITHOUT reading any skill files:
- Task: <eval prompt>
- Save outputs to: <workspace>/iteration-N/eval-<name>/without_skill/
Show full SKILL.md (634 more words)Show less
While runs execute, draft assertions

Use the time productively. Draft quantitative assertions for each test case — objectively verifiable checks with descriptive names. Explain to the user what the assertions check.

Good assertions: "output_file_exists", "contains_required_sections", "no_placeholder_text", "follows_naming_convention" Bad assertions: "output_is_good", "looks_right" (subjective — evaluate qualitatively instead)

Step 6: Evaluate Results

When runs complete:

  1. Grade each run — check assertions programmatically where possible (write and run a script), use judgment for the rest
  2. Present to user — show each test case's prompt, skill output, and baseline output side by side. Ask for feedback on each.
  3. Analyze patterns — look for assertions that always pass regardless of skill (non-discriminating), high-variance results (flaky), and time/token tradeoffs

Empty feedback from the user means they thought it was fine. Focus improvements on test cases with specific complaints.

Step 7: Improve the Skill

This is the heart of the loop. Principles for making improvements:

Generalize from feedback. The skill will be used across many prompts — if a fix only works for the specific test case, it's overfitting. Rather than fiddly overfit changes or oppressive MUSTs, try different metaphors or recommend different working patterns.

Keep the prompt lean. Remove things that aren't pulling their weight. Read the subagent transcripts, not just final outputs — if the skill makes the model waste time on unproductive steps, cut those instructions.

Explain the why. Try to transmit understanding, not just rules. The model has good theory of mind and can go beyond rote instructions when it understands the reasoning.

Look for repeated work. If all test runs independently wrote similar helper scripts or took the same multi-step approach, that's a signal to bundle that script in scripts/.

Draft, then review. Write a revision, then look at it with fresh eyes and improve before committing.

After improving: rerun all test cases into a new iteration-<N+1>/ directory, present results with previous iteration for comparison, get feedback, repeat.

Step 8: Validate

Run the validator before finalizing:

bash
python ${CLAUDE_SKILL_DIR}/scripts/validate_skill.py .claude/skills/<skill-name>

This checks structure, frontmatter, naming conventions, placeholder detection, resource references, and description quality.

Step 9: Optimize Description (Optional)

After the skill is working well, offer to optimize the description for better triggering accuracy.

Generate trigger eval queries

Create 20 eval queries — a mix of should-trigger (8-10) and should-not-trigger (8-10).

Should-trigger queries: Different phrasings of the same intent — formal and casual. Include cases where the user doesn't name the skill explicitly but clearly needs it. Cover uncommon use cases and cases where this skill competes with another.

Should-not-trigger queries: Focus on near-misses — queries that share keywords but need something different. Adjacent domains, ambiguous phrasing where keyword matching would trigger but shouldn't. Avoid obviously irrelevant queries ("write a fibonacci function" as negative for a PDF skill tests nothing).

Queries must be realistic — concrete, specific, with detail. Not "Format this data" but "ok so my boss sent me this xlsx and she wants me to add a profit margin column, revenue is in C and costs in D".

Test triggering

For each query, test whether Claude would trigger the skill by examining the description match. Adjust the description to improve accuracy on the eval set, being careful not to overfit.

Apply

Update the SKILL.md frontmatter with the optimized description. Show the user before/after.

Anti-patterns

  • Testing your own skill — if you wrote the skill and you're also running it with full context, the test is compromised. Use separate subagents that don't share context with the drafting process.
  • Skipping baseline runs — without a baseline, you can't tell if the skill actually helped or if Claude would have done fine without it.
  • Overfitting to test cases — if the skill only works for 3 specific prompts, it's useless at scale. Generalize.
  • MUST/NEVER overload — rigid constraints are a code smell in skills. Explain reasoning instead.
  • Skipping user review — always get human feedback on outputs before revising. The user sees things you can't.

© deepklarity, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (scripts, references) in .claude/skills/hk-skill-creator of deepklarity/harness-kit.

  • SKILL.md
  • references/writing_guide.md
  • scripts/init_skill.py
  • scripts/validate_skill.py

Open the folder on GitHubat commit 87305cd

Compare with similar skills

Hk Skill Creator next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Hk Skill Creator compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Hk Skill Creator this skilldeepklarity/harness-kit100—~3.2kAutomated safety check: NotesMIT
Skill CreatorAzure/azqr79689 repos~8.2kAutomated safety check: PassApache-2.0
Skillforgetripleyak/SkillForge906—~2.3kAutomated safety check: NotesMIT
Zach Seller Skill Creatorzach22-1999/amazon-skills2091 repos~3.9kAutomated safety check: PassApache-2.0
Skill CreatorAgentTeam-TaichuAI/ScienceClaw671—~10kAutomated safety check: PassApache-2.0
Skill Creatorluongnv89/asm954—~5.3kAutomated safety check: PassMIT

Similar skills

  • Skill Creator

    Azure/azqr

    Official

    Create new skills, modify and improve existing skills, and measure skill performance.

    796 GitHub starsUsed in 89 repos~8.2k tokens
    Agent WorkflowsAuto-check passed
  • Skillforge

    tripleyak/SkillForge

    A skill your agent uses when creating, improving, finding, or auditing agent skills - the user says 'create a skill', 'do I have a skill for X', 'improve the X skill', 'which skill should I use'…

    906 GitHub stars~2.3k tokensUpdated 2 mo ago
    Agent WorkflowsAuto-check: notes
  • Zach Seller Skill Creator

    zach22-1999/amazon-skills

    亚马逊卖家专用的 skill 创建器(中文)。当用户想把一个亚马逊运营/自媒体/日常工作流程变成可复用的 skill 时使用。触发场景包括但不限于:用户说"我想做一个 skill""把这个流程变成 skill""帮我写个自动化""优化我已有的 skill""给这个工作流做个自动化",即使用户没用"skill"这个词,只要在描述"以后每次都这样做"的重复性工作时也应触发。本 skill…

    209 GitHub starsUsed in 1 repo~3.9k tokens
    Agent WorkflowsAuto-check passed
  • Skill Creator

    AgentTeam-TaichuAI/ScienceClaw

    Create new skills, modify and improve existing skills, and measure skill performance.

    671 GitHub stars~10k tokensUpdated 5 mo ago
    Agent WorkflowsAuto-check passed
  • Skill Creator

    luongnv89/asm

    Create a skill or bring an existing one up to the same standard (validate + asm eval fix loop); run evals, tune triggering.

    954 GitHub stars~5.3k tokensUpdated 4 days ago
    Agent WorkflowsAuto-check passed
  • Skill Creator

    feiskyer/claude-code-settings

    Create, refine, and benchmark agent skills. An agent skill from feiskyer/claude-code-settings.

    1.7k GitHub stars~7.6k tokensUpdated 12 days ago
    Agent WorkflowsAuto-check passed

More from deepklarity/harness-kit

All 18 skills in this repo
  • Hk Arch Audit

    deepklarity/harness-kit

    Run comprehensive agent-native architecture review with scored principles.

    100 GitHub stars~1.1k tokensUpdated 2 mo ago
    Auto-check: notes
  • Hk Mock First

    deepklarity/harness-kit

    Mock-first, layer-by-layer feature development. An agent skill from deepklarity/harness-kit.

    100 GitHub stars~3.9k tokensUpdated 2 mo ago
    Auto-check: notes
  • Hk Autonomy Audit

    deepklarity/harness-kit

    Audit whether an AI agent can autonomously close the loop on problems in a given area — from discovering a symptom to verifying a fix — without human intervention.

    100 GitHub stars~2.5k tokensUpdated 2 mo ago
    Auto-check: notes
  • Hk Breadcrumb Creator

    deepklarity/harness-kit

    Traces a workflow end-to-end through the harness-kit monorepo and creates a breadcrumb analysis doc in docs/breadcrumbanalysis/.

    100 GitHub stars~3k tokensUpdated 2 mo ago
    Auto-check: notes
  • Hk Changelog

    deepklarity/harness-kit

    Generate changelog entries from git diffs, prepend to CHANGELOG.md, and optionally commit + PR.

    100 GitHub stars~1.6k tokensUpdated 2 mo ago
    Auto-check: notes
  • Hk Compound

    deepklarity/harness-kit

    Compound a learning into a reusable pattern. An agent skill from deepklarity/harness-kit.

    100 GitHub stars~1.7k tokensUpdated 2 mo ago
    Auto-check: notes

Categories

Questions about Hk Skill Creator

What does Hk Skill Creator do?

Create new skills, modify and improve existing skills, and measure skill performance. Hk Skill Creator is an agent skill from deepklarity/harness-kit. Create new skills, modify and improve existing skills, and measure skill performance.

When should I use Hk Skill Creator?

Hk Skill Creator fits situations like: users want to create a skill from scratch; optimize an existing skill; run evals to test a skill; benchmark skill performance with variance analysis.

How do I install Hk Skill Creator in Claude Code?

Run `npx skills add deepklarity/harness-kit --skill hk-skill-creator -a claude-code`. Or copy the skill folder (.claude/skills/hk-skill-creator in deepklarity/harness-kit) into .claude/skills/hk-skill-creator in your project. Claude Code loads it when a task matches its description.

How do I install Hk Skill Creator in Codex?

Run `npx skills add deepklarity/harness-kit --skill hk-skill-creator -a codex`. Or copy the skill folder (.claude/skills/hk-skill-creator in deepklarity/harness-kit) into .agents/skills/hk-skill-creator in your project. Codex loads it when a task matches its description.

Can I use Hk Skill Creator in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add deepklarity/harness-kit --skill hk-skill-creator -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/hk-skill-creator, .gemini/skills/hk-skill-creator, .github/skills/hk-skill-creator and .opencode/skills/hk-skill-creator in your project.

What does Hk Skill Creator need to run?

Going by SKILL.md and its folder, Hk Skill Creator needs Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python 3. Its frontmatter pre-approves these tools: Bash, Read, Edit, Write, Grep, Glob, Agent, Task, TaskCreate, TaskUpdate.

Does Hk Skill Creator access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Hk Skill Creator safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Hk Skill Creator use?

Hk Skill Creator is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Hk Skill Creator use?

About 3.2k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.3k tokens, read only when the agent opens those files.

What are the alternatives to Hk Skill Creator?

Skills that share tags, products or a category with Hk Skill Creator: Skill Creator (Azure/azqr, 796 stars), Skillforge (tripleyak/SkillForge, 906 stars), Zach Seller Skill Creator (zach22-1999/amazon-skills, 209 stars) and Skill Creator (AgentTeam-TaichuAI/ScienceClaw, 671 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Hk Skill Creator?

deepklarity (a GitHub organization) maintains it in deepklarity/harness-kit, which has 100 GitHub stars. The repository holds 18 skills in this directory. The repository was last updated on July 15, 2026.

Source: deepklarity/harness-kit on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.