Skill Creator
ZS520L/HanakoPro
Create new skills, modify and improve existing skills, and measure skill performance.
Evaluates agent skills with a Go CLI that runs YAML-defined benchmarks, compares runs and scores the quality of SKILL.md frontmatter.
$ npx skills add microsoft/waza --skill waza -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install microsoft/waza waza --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/microsoft/waza.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/waza .claude/skills/waza && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "waza" agent skill from https://github.com/microsoft/waza/tree/main/skills/waza into .claude/skills/waza/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "waza", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/microsoft/waza/tree/main/skills/wazaType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add microsoft/waza --skill waza -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install microsoft/waza waza --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/microsoft/waza.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/waza .agents/skills/waza && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "waza" agent skill from https://github.com/microsoft/waza/tree/main/skills/waza into .agents/skills/waza/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "waza", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add microsoft/waza --skill waza -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install microsoft/waza waza --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/microsoft/waza.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/waza .cursor/skills/waza && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "waza" agent skill from https://github.com/microsoft/waza/tree/main/skills/waza into .cursor/skills/waza/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "waza", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/microsoft/waza.git --path skills/waza--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add microsoft/waza --skill waza -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install microsoft/waza waza --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/microsoft/waza.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/waza .gemini/skills/waza && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "waza" agent skill from https://github.com/microsoft/waza/tree/main/skills/waza into .gemini/skills/waza/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "waza", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install microsoft/waza wazaInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add microsoft/waza --skill waza -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/microsoft/waza.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/waza .github/skills/waza && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "waza" agent skill from https://github.com/microsoft/waza/tree/main/skills/waza into .github/skills/waza/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "waza", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add microsoft/waza --skill waza -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install microsoft/waza waza --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/microsoft/waza.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/waza .opencode/skills/waza && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "waza" agent skill from https://github.com/microsoft/waza/tree/main/skills/waza into .opencode/skills/waza/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "waza", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
wazaEvaluates agent skills with a Go CLI that runs YAML-defined benchmarks, compares runs and scores the quality of SKILL.md frontmatter.
waza is a Go command-line tool for testing agent skills. You describe test cases in a YAML eval spec, run them with waza run against an agent engine (the mock engine by default, with fixtures supplied through --context-dir), and score the results with pluggable validators, which the description lists as code and regex checks.
Other commands round out the loop. waza init creates eval.yaml with an example task and fixture, waza generate builds an eval suite from an existing SKILL.md by reading its name and description, and waza compare lines up several result files to show per-task score deltas, pass-rate differences and aggregate statistics.
waza dev scores a skill's frontmatter on four compliance levels from Low to High. It checks that the description is at least 150 characters (1024 at most), has trigger and anti-trigger phrases and routing markers such as INVOKES, and stays within a token budget with a soft limit of 500 and a hard limit of 5000. Creating skills from scratch and token counting are left to other tools.
Read from SKILL.md and the folder at commit 774df00. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are bash and yaml).
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Waza Skill Evaluator loads about 2k tokens when it runs. Until then it costs about 161 tokens; SKILL.md has 307 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from microsoft/waza at commit 774df00, republished under its MIT licence (© microsoft). 307 words, ~1,953 tokens.
.claude/skills/waza/SKILL.md (or your agent's skills folder)."The way of technique — measure, refine, master."
A Go CLI tool for evaluating AI agent skills through structured benchmarks. Define test cases in YAML, run them against agent engines, and validate results with pluggable scoring validators.
When user says "waza help" or asks how to use waza:
╔══════════════════════════════════════════════════════════════════╗
║ WAZA - CLI Tool for Evaluating Agent Skills ║
╠══════════════════════════════════════════════════════════════════╣
║ ║
║ COMMANDS: ║
║ waza run <eval.yaml> # Run an evaluation benchmark ║
║ waza init [directory] # Initialize a new eval suite ║
║ waza generate <SKILL.md> # Generate eval from SKILL.md ║
║ waza compare <r1> <r2> ... # Compare result files ║
║ waza dev [skill-path] # Improve SKILL.md compliance ║
║ ║
║ RUN FLAGS: ║
║ --context-dir, -c Fixtures directory (default: ./fixtures) ║
║ --output, -o Save results JSON to file ║
║ --verbose, -v Verbose output ║
║ --task, -t Filter tasks by name (repeatable) ║
║ --parallel, -p Run tasks in parallel ║
║ --workers, -w Number of parallel workers ║
║ --transcript-dir Save per-task transcripts ║
║ ║
║ COMPARE FLAGS: ║
║ --format, -f Output format: table or json ║
║ ║
║ GENERATE FLAGS: ║
║ --output-dir, -d Output directory for generated files ║
║ ║
║ DEV FLAGS: ║
║ --target Adherence level: low|medium|high ║
║ --max-iterations Max improvement iterations (default: 5) ║
║ --auto Auto-apply without prompting ║
║ ║
║ WORKFLOW: ║
║ 1. waza init my-eval # Scaffold eval suite ║
║ 2. Edit eval.yaml + tasks # Define test cases ║
║ 3. waza run eval.yaml -v # Execute benchmark ║
║ 4. waza compare a.json b.json # Cross-model comparison ║
║ ║
║ FIXTURE ISOLATION: ║
║ Each task gets a fresh temp workspace with fixtures copied ║
║ in. Original fixtures are never modified. ║
║ ║
╚══════════════════════════════════════════════════════════════════╝waza runRun an evaluation benchmark from a YAML spec file.
# Run with default mock engine
waza run path/to/eval.yaml --context-dir path/to/fixtures
# Verbose output with results saved
waza run eval.yaml -c ./fixtures -v -o results.json
# Filter to specific tasks
waza run eval.yaml -t "task-name-1" -t "task-name-2"
# Parallel execution
waza run eval.yaml --parallel --workers 4
# Save per-task transcripts
waza run eval.yaml --transcript-dir ./transcriptswaza initInitialize a new evaluation suite with a compliant directory structure.
# Initialize in current directory
waza init
# Initialize in a named directory
waza init my-eval-suiteCreates: eval.yaml, tasks/ with example task, fixtures/ with example fixture.
waza generateGenerate an eval suite from an existing SKILL.md file.
# Generate eval from SKILL.md
waza generate path/to/SKILL.md
# Specify output directory
waza generate SKILL.md --output-dir ./my-evalParses YAML frontmatter (name, description) and creates eval.yaml, starter tasks, and fixtures.
waza compareCompare results from multiple evaluation runs side by side.
# Compare two result files
waza compare run1.json run2.json
# Compare three or more
waza compare gpt4.json claude.json gemini.json
# JSON output
waza compare run1.json run2.json --format jsonShows per-task score deltas, pass rate differences, and aggregate statistics.
waza devIteratively improve SKILL.md frontmatter compliance with automated scoring.
# Score current skill and suggest improvements
waza dev skills/my-skill
# Target high compliance level
waza dev skills/my-skill --target high
# Auto-apply improvements without prompts
waza dev skills/my-skill --target medium --auto --max-iterations 3Compliance Levels:
Scoring Checks:
Coming Soon: Trigger accuracy tests (#36), --skip-integration (#37), --fast (#38), improvement suggestions engine (#34).
name: my-eval
skill: my-skill
version: "1.0"
executor: mock # or copilot-sdk
tasks:
- id: task-1
name: "Describe the task"
prompt: "Your prompt to the agent"
expected: "Expected behavior"
validators:
- type: code
config:
language: go
- type: text
config:
pattern: "expected pattern"| Engine | Use | Description |
|---|---|---|
mock | Testing | Returns canned responses for validator development |
copilot-sdk | Production | Executes via Copilot CLI SDK |
| Validator | What it checks |
|---|---|
code | Code compiles / passes syntax check |
regex | Output matches regex pattern |
| Setting | Flag | Default |
|---|---|---|
| Fixtures dir | --context-dir | ./fixtures |
| Output file | --output | (none) |
| Verbose | --verbose | false |
| Parallel | --parallel | false |
| Workers | --workers | CPU count |
| Transcript dir | --transcript-dir | (none) |
Each task produces an EvaluationOutcome with:
| Field | Description |
|---|---|
score | 0.0–1.0 normalized score |
pass | Boolean pass/fail |
validator_results | Per-validator details |
duration | Execution time |
© microsoft, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/waza of microsoft/waza.
Open the folder on GitHubat commit 774df00
Waza Skill Evaluator next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Waza Skill Evaluator this skillmicrosoft/waza | 1.4k | — | ~2k | Automated safety check: Pass | MIT | |
| Skill CreatorZS520L/HanakoPro | 102 | — | ~7.4k | Automated safety check: Pass | Apache-2.0 | |
| Skill Creatorhimself65/finance-skills | 3.4k | — | ~3.8k | Automated safety check: Pass | MIT | |
| Skill Eval ImproveArenukvern/mcp_flutter | 386 | — | ~2.4k | Automated safety check: Pass | MIT | |
| Skill CreatorAzure/azqr | 795 | 89 repos | ~8.2k | Automated safety check: Pass | Apache-2.0 | |
| Darwin Skill Optimizeralchaincyf/darwin-skill | 6.2k | 1 repos | ~4.7k | Automated safety check: Pass | MIT |
ZS520L/HanakoPro
Create new skills, modify and improve existing skills, and measure skill performance.
himself65/finance-skills
Create, improve, and evaluate agent skills (SKILL.md plus reference files).
Arenukvern/mcp_flutter
Improves Agent Skills via validate → rule-based eval cases → plugin-eval → prompt evals → bounded edits with held-out gates.
Azure/azqr
Create new skills, modify and improve existing skills, and measure skill performance.
alchaincyf/darwin-skill
Scores SKILL.md files on a nine-dimension rubric, then improves them in a keep-or-revert loop with independent judge agents, test prompts, git history and human checkpoints.
rohitg00/ai-engineering-from-scratch
Evaluates an Agent Skill bundle before release for structure, trigger quality, artifact improvement, script correctness, safety, installed-tree integrity and host portability.
microsoft/waza
Shows a categorized, interactive menu of common Squad operations, such as install, upgrade and team management, and collects arguments before running anything.
microsoft/waza
Shared collaboration rules for a team of squad agents covering worktree awareness, writing decisions to an inbox, cross-agent requests and reviewer lockout.
microsoft/waza
Walks through releasing a new version of the waza azd extension: changelog from commits, semver bump with your confirmation, and a release PR.
microsoft/waza
Dev-first branching model for the Squad project: feature work branches from dev, issue branches follow a naming rule and parallel issues use git worktrees.
microsoft/waza
Reviewer rejection workflow and strict lockout semantics. An agent skill from microsoft/waza.
microsoft/waza
Walks you through creating, running and reading waza evals for an agent skill, then proposes concrete fixes when tasks fail or the score is low.
Categories
Evaluates agent skills with a Go CLI that runs YAML-defined benchmarks, compares runs and scores the quality of SKILL.md frontmatter. waza is a Go command-line tool for testing agent skills. You describe test cases in a YAML eval spec, run them with waza run against an agent engine (the mock engine by default, with fixtures supplied through --context-dir), and score the results with pluggable validators, which the description lists as code and regex checks.
Waza Skill Evaluator fits situations like: benchmarking an agent skill against a set of YAML test cases; generating a starter eval suite from an existing SKILL.md; comparing two evaluation runs to see which tasks improved; scoring a skill's frontmatter and description for trigger coverage.
Run `npx skills add microsoft/waza --skill waza -a claude-code`. Or copy the skill folder (skills/waza in microsoft/waza) into .claude/skills/waza in your project. Claude Code loads it when a task matches its description.
Run `npx skills add microsoft/waza --skill waza -a codex`. Or copy the skill folder (skills/waza in microsoft/waza) into .agents/skills/waza in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add microsoft/waza --skill waza -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/waza, .gemini/skills/waza, .github/skills/waza and .opencode/skills/waza in your project.
SKILL.md names no scripts, command-line tools or credentials: Waza Skill Evaluator is instructions for the agent only. Our summary lists: The waza command-line tool; An agent engine such as the Copilot SDK executor, or the built-in mock engine for dry runs.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Waza Skill Evaluator is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2k tokens (SKILL.md is roughly 7.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Waza Skill Evaluator: Skill Creator (ZS520L/HanakoPro, 102 stars), Skill Creator (himself65/finance-skills, 3.4k stars), Skill Eval Improve (Arenukvern/mcp_flutter, 386 stars) and Skill Creator (Azure/azqr, 795 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
microsoft (a GitHub organization, an official publisher) maintains it in microsoft/waza, which has 1,403 GitHub stars. The repository holds 16 skills in this directory. The repository was last updated on October 6, 2026.
Source: microsoft/waza on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.