Agent skill

Empirical Prompt Tuning

by mizchi in mizchi/skills

Methodology for iteratively improving agent-facing instructions (skills / slash commands / CLAUDE.md / code-gen prompts) via bias-free executor + two-sided evaluation (self-report + instruction-side…

No licenceAuto-check passedAgent Workflows

Install Empirical Prompt Tuning

skills CLI
$ npx skills add mizchi/skills --skill empirical-prompt-tuning -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install mizchi/skills empirical-prompt-tuning --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/mizchi/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/empirical-prompt-tuning .claude/skills/empirical-prompt-tuning && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
empirical-prompt-tuning
GitHub stars
360
Token cost
~5k tokens
SKILL.md length
2,315 words
Files
3
Skills in repo
69
Repo updated
First seen
Licence
None found

At a glance

Methodology for iteratively improving agent-facing instructions (skills / slash commands / CLAUDE.md / code-gen prompts) via bias-free executor + two-sided evaluation (self-report + instruction-side…

  • Works in 8 steps: Iteration 0 — description / body… → Baseline preparation: Fix the target… → Bias-free read: Have a "blank-slate"… → …
  • Explicitly asks for an empirical eval of a prompt
  • SKILL.md covers When to use, Workflow, Evaluation axes and Subagent invocation contract, plus 8 more sections
  • Calls npx

What it does

Empirical Prompt Tuning is an agent skill from mizchi/skills. Methodology for iteratively improving agent-facing instructions (skills / slash commands / CLAUDE.md / code-gen prompts) via bias-free executor + two-sided evaluation (self-report + instruction-side metrics). Meta-skill, invoke ONLY when the user explicitly asks for an "empirical" eval of a prompt or skill, or for the Iter-0 description / body consistency check. Do NOT auto-invoke after every skill edit; this loop is operator-triggered by name.

Its SKILL.md is about 5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `README.md` and `SKILL-ja.md`).

It sits in Agent Workflows, covering Hooks and plugins and Agent instruction files. The repository describes itself as: Agent skills by mizchi, distributed via APM.

When your agent uses it

  • Explicitly asks for an empirical eval of a prompt
  • For the Iter-0 description / body consistency check

Example prompts

  • “empirical”
  • “/empirical-prompt-tuning”

Requirements

  • Node.js

Workflow steps

8 steps, taken from the first numbered list in SKILL.md.

  1. Iteration 0 — description / body consistency check (static, no dispatch needed)
  2. Baseline preparation: Fix the target prompt and prepare the following two things.
  3. Bias-free read: Have a "blank-slate" executor read the instruction. Dispatch a new subagent via the Task tool. Do not substitute with a…
  4. Execution: Hand the subagent a prompt that follows the subagent invocation contract described below, and have it execute the scenario. The…
  5. Two-sided evaluation: Record the following from the returned results.
  6. Apply the diff: Put the minimum fix into the prompt to eliminate the unclear points. One theme per iteration (multiple related fixes are…
  7. Re-evaluate: Run 2 → 5 again with a new subagent (do not reuse the same agent: it has learned the previous improvements). Increase…
  8. Convergence check: The rough rule is "stop when 2 consecutive iterations have zero new unclear points AND metric improvements fall below…

What it can do on your machine

Read from SKILL.md and the folder at commit 62f5808. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • npx

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use npx, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Empirical Prompt Tuning loads about 5k tokens when it runs. Until then it costs about 118 tokens; SKILL.md has 2,315 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~118
When it runs · the whole SKILL.md, loaded when a task matches
~5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

Without a licence we can't republish the file, so here is its outline and opening line. It has 2,315 words (~4,956 tokens).

“The author of a prompt cannot judge its quality. The clearer the writer thinks something is, the more likely another agent will stumble on it. The core of this skill is to have a bias-free executor actually run the instruction…”

— opening of SKILL.md by mizchi
name
empirical-prompt-tuning

Read the full SKILL.md on GitHub

Files

SKILL.md and 2 other files in empirical-prompt-tuning of mizchi/skills.

  • SKILL.md
  • README.md
  • SKILL-ja.md

Open the folder on GitHubat commit 62f5808

Compare with similar skills

Empirical Prompt Tuning next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Empirical Prompt Tuning compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Empirical Prompt Tuning this skillmizchi/skills360—~5kAutomated safety check: PassNone
Agent Setup Health Audittw93/Waza7.2k—~5.2kAutomated safety check: NotesMIT
Working With Claude Code Docsobra/superpowers-developing-for-claude-code142—~1.5kAutomated safety check: PassNone
Directional Promptingkingbootoshi/directional-prompting143—~2.3kAutomated safety check: PassMIT
Agenticashibing624/agentica352—~1.8kAutomated safety check: NotesApache-2.0
Claude Code Mastery Squadohmyjahh/xquads-squads277—~1.1kAutomated safety check: PassMIT

Similar skills

  • Audits a project's agent configuration, instruction drift, hooks, MCP and AI maintainability, then reports prioritized findings with evidence and next actions.

    7.2k GitHub stars~5.2k tokensUpdated today
    Agent WorkflowsAuto-check: notes
  • Working With Claude Code Docs

    obra/superpowers-developing-for-claude-code

    Looks up official Claude Code documentation stored as reference files instead of guessing about CLI commands, configuration, or plugin APIs.

    142 GitHub stars~1.5k tokensUpdated 10 mo ago
    Agent WorkflowsAuto-check passed
  • Directional Prompting

    kingbootoshi/directional-prompting

    Write prompts, system instructions, agent directives, slash commands, and skill descriptions using two stacked layers — outcome-first (define the destination, success criteria, stopping condition)…

    143 GitHub stars~2.3k tokensUpdated 4 mo ago
    Agent WorkflowsAuto-check passed
  • Agentica

    shibing624/agentica

    How to answer questions about the agentica product you are running inside — CLI flags, config.yaml profiles, API keys, models, sessions, resume, workspace, AGENTS.md standing rules, skills, logs…

    352 GitHub stars~1.8k tokensUpdated yesterday
    Agent WorkflowsAuto-check: notes
  • Claude Code Mastery Squad

    ohmyjahh/xquads-squads

    Routes a request to one of eight specialist agents covering hooks, skills, subagents, MCP integration and context engineering for Claude Code.

    277 GitHub stars~1.1k tokensUpdated 12 days ago
    Agent WorkflowsAuto-check passed
  • Vault Setup

    earlyaidopters/second-brain

    Interactive Obsidian vault configurator. An agent skill from earlyaidopters/second-brain.

    194 GitHub stars~1.4k tokensUpdated 6 mo ago
    Agent WorkflowsAuto-check passed

More from mizchi/skills

All 69 skills in this repo
  • Publish Agent Skill

    mizchi/skills

    Package, publish, verify and update an agent skill through a Claude Code plugin marketplace, APM or the npx skills CLI.

    360 GitHub stars~1.7k tokensUpdated 9 days ago
    Auto-check passed
  • AI Index

    mizchi/skills

    Method and tooling for measuring how AI-generated a piece of prose reads, in Japanese or English.

    360 GitHub stars~3.6k tokensUpdated 9 days ago
    Auto-check passed
  • Cloudflare Deploy

    mizchi/skills

    Deploy applications and infrastructure to Cloudflare with the cf CLI and typed cloudflare.config.ts.

    360 GitHub stars~2.9k tokensUpdated 9 days ago
    Auto-check: notes
  • Review Image

    mizchi/skills

    Review screenshots or other images with OpenRouter vision models via bundled Deno scripts.

    360 GitHub stars~1.3k tokensUpdated 9 days ago
    Auto-check passed
  • Post-generation safety checks for sqlc-gen-moonbit + Cloudflare D1.

    360 GitHub stars~995 tokensUpdated 9 days ago
    Auto-check passed
  • Apm Usage

    mizchi/skills

    Reference for APM (Agent Package Manager) — apm.yml syntax, install / uninstall / update commands, target detection, lockfile workflow.

    360 GitHub stars~1.9k tokensUpdated 9 days ago
    Auto-check passed

Categories

Questions about Empirical Prompt Tuning

What does Empirical Prompt Tuning do?

Methodology for iteratively improving agent-facing instructions (skills / slash commands / CLAUDE.md / code-gen prompts) via bias-free executor + two-sided evaluation (self-report + instruction-side…. Empirical Prompt Tuning is an agent skill from mizchi/skills.md / code-gen prompts) via bias-free executor + two-sided evaluation (self-report + instruction-side metrics).

When should I use Empirical Prompt Tuning?

Empirical Prompt Tuning fits situations like: explicitly asks for an empirical eval of a prompt; for the Iter-0 description / body consistency check.

How do I install Empirical Prompt Tuning in Claude Code?

Run `npx skills add mizchi/skills --skill empirical-prompt-tuning -a claude-code`. Or copy the skill folder (empirical-prompt-tuning in mizchi/skills) into .claude/skills/empirical-prompt-tuning in your project. Claude Code loads it when a task matches its description.

How do I install Empirical Prompt Tuning in Codex?

Run `npx skills add mizchi/skills --skill empirical-prompt-tuning -a codex`. Or copy the skill folder (empirical-prompt-tuning in mizchi/skills) into .agents/skills/empirical-prompt-tuning in your project. Codex loads it when a task matches its description.

Can I use Empirical Prompt Tuning in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mizchi/skills --skill empirical-prompt-tuning -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/empirical-prompt-tuning, .gemini/skills/empirical-prompt-tuning, .github/skills/empirical-prompt-tuning and .opencode/skills/empirical-prompt-tuning in your project.

What does Empirical Prompt Tuning need to run?

Going by SKILL.md and its folder, Empirical Prompt Tuning needs the command-line tools its instructions call (npx). Our summary lists: Node.js.

Does Empirical Prompt Tuning access the network?

SKILL.md contains no URLs. Its commands use npx, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Empirical Prompt Tuning safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Empirical Prompt Tuning use?

No licence was found for Empirical Prompt Tuning or its repository. Without one, default copyright applies: ask the author before reusing or redistributing it.

How many tokens does Empirical Prompt Tuning use?

About 5k tokens (SKILL.md is roughly 20k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Empirical Prompt Tuning?

Skills that share tags, products or a category with Empirical Prompt Tuning: Agent Setup Health Audit (tw93/Waza, 7.2k stars), Working With Claude Code Docs (obra/superpowers-developing-for-claude-code, 142 stars), Directional Prompting (kingbootoshi/directional-prompting, 143 stars) and Agentica (shibing624/agentica, 352 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Empirical Prompt Tuning?

mizchi (a GitHub user) maintains it in mizchi/skills, which has 360 GitHub stars. The repository holds 69 skills in this directory. The repository was last updated on October 2, 2026.

Source: mizchi/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.