Official agent skill

Waza Skill Evaluator

by microsoft in microsoft/waza

Evaluates agent skills with a Go CLI that runs YAML-defined benchmarks, compares runs and scores the quality of SKILL.md frontmatter.

OfficialMITAuto-check passedAgent Workflows

Install Waza Skill Evaluator

skills CLI
$ npx skills add microsoft/waza --skill waza -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install microsoft/waza waza --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/microsoft/waza.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/waza .claude/skills/waza && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
waza
GitHub stars
1.4k
Token cost
~2k tokens
SKILL.md length
307 words
Files
1
Skills in repo
16
Repo updated
First seen
Licence
MIT

At a glance

Evaluates agent skills with a Go CLI that runs YAML-defined benchmarks, compares runs and scores the quality of SKILL.md frontmatter.

  • Benchmarking an agent skill against a set of YAML test cases
  • SKILL.md covers Help, Commands, Evaluation Spec Format and Engines, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Generating a starter eval suite from an existing SKILL.md

What it does

waza is a Go command-line tool for testing agent skills. You describe test cases in a YAML eval spec, run them with waza run against an agent engine (the mock engine by default, with fixtures supplied through --context-dir), and score the results with pluggable validators, which the description lists as code and regex checks.

Other commands round out the loop. waza init creates eval.yaml with an example task and fixture, waza generate builds an eval suite from an existing SKILL.md by reading its name and description, and waza compare lines up several result files to show per-task score deltas, pass-rate differences and aggregate statistics.

waza dev scores a skill's frontmatter on four compliance levels from Low to High. It checks that the description is at least 150 characters (1024 at most), has trigger and anti-trigger phrases and routing markers such as INVOKES, and stays within a token budget with a soft limit of 500 and a hard limit of 5000. Creating skills from scratch and token counting are left to other tools.

When your agent uses it

  • Benchmarking an agent skill against a set of YAML test cases
  • Generating a starter eval suite from an existing SKILL.md
  • Comparing two evaluation runs to see which tasks improved
  • Scoring a skill's frontmatter and description for trigger coverage

Example prompts

  • “Run the waza eval in evals/pdf-forms/eval.yaml and summarize which tasks fail.”
  • “Generate an eval suite for skills/my-skill/SKILL.md.”
  • “Compare run1.json and run2.json and tell me which tasks regressed.”
  • “Run waza dev on my skill and tell me what is missing from its description.”

Requirements

  • The waza command-line tool
  • An agent engine such as the Copilot SDK executor, or the built-in mock engine for dry runs

What it can do on your machine

Read from SKILL.md and the folder at commit 774df00. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are bash and yaml).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Waza Skill Evaluator loads about 2k tokens when it runs. Until then it costs about 161 tokens; SKILL.md has 307 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~161
When it runs · the whole SKILL.md, loaded when a task matches
~2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from microsoft/waza at commit 774df00, republished under its MIT licence (© microsoft). 307 words, ~1,953 tokens.

Download SKILL.mdSave it as .claude/skills/waza/SKILL.md (or your agent's skills folder).
name
waza
description
**WORKFLOW SKILL** - Evaluate AI agent skills using structured benchmarks with YAML specs, fixture isolation, and pluggable validators. USE FOR: run waza, waza help, run eval, run benchmark, evaluate skill, test agent, generate eval suite, init eval, compare results, score agent, agent evaluation, skill testing, cross-model comparison. DO NOT USE FOR: improving skill frontmatter (use waza dev), creating new skills from scratch (use skill-creator), token counting or budget checks (use waza tokens). INVOKES: Copilot SDK executor, mock engine, code/regex validators. FOR SINGLE OPERATIONS: use waza run directly for a single benchmark.

Waza

"The way of technique — measure, refine, master."

A Go CLI tool for evaluating AI agent skills through structured benchmarks. Define test cases in YAML, run them against agent engines, and validate results with pluggable scoring validators.

Help

When user says "waza help" or asks how to use waza:

╔══════════════════════════════════════════════════════════════════╗
║  WAZA - CLI Tool for Evaluating Agent Skills                     ║
╠══════════════════════════════════════════════════════════════════╣
║                                                                  ║
║  COMMANDS:                                                       ║
║    waza run <eval.yaml>        # Run an evaluation benchmark     ║
║    waza init [directory]       # Initialize a new eval suite     ║
║    waza generate <SKILL.md>    # Generate eval from SKILL.md     ║
║    waza compare <r1> <r2> ...  # Compare result files            ║
║    waza dev [skill-path]       # Improve SKILL.md compliance     ║
║                                                                  ║
║  RUN FLAGS:                                                      ║
║    --context-dir, -c   Fixtures directory (default: ./fixtures)  ║
║    --output, -o        Save results JSON to file                 ║
║    --verbose, -v       Verbose output                            ║
║    --task, -t          Filter tasks by name (repeatable)         ║
║    --parallel, -p      Run tasks in parallel                     ║
║    --workers, -w       Number of parallel workers                ║
║    --transcript-dir    Save per-task transcripts                 ║
║                                                                  ║
║  COMPARE FLAGS:                                                  ║
║    --format, -f        Output format: table or json              ║
║                                                                  ║
║  GENERATE FLAGS:                                                 ║
║    --output-dir, -d    Output directory for generated files      ║
║                                                                  ║
║  DEV FLAGS:                                                      ║
║    --target            Adherence level: low|medium|high          ║
║    --max-iterations    Max improvement iterations (default: 5)   ║
║    --auto              Auto-apply without prompting              ║
║                                                                  ║
║  WORKFLOW:                                                       ║
║    1. waza init my-eval        # Scaffold eval suite             ║
║    2. Edit eval.yaml + tasks   # Define test cases               ║
║    3. waza run eval.yaml -v    # Execute benchmark               ║
║    4. waza compare a.json b.json  # Cross-model comparison       ║
║                                                                  ║
║  FIXTURE ISOLATION:                                              ║
║    Each task gets a fresh temp workspace with fixtures copied    ║
║    in. Original fixtures are never modified.                     ║
║                                                                  ║
╚══════════════════════════════════════════════════════════════════╝

Commands

waza run

Run an evaluation benchmark from a YAML spec file.

bash
# Run with default mock engine
waza run path/to/eval.yaml --context-dir path/to/fixtures

# Verbose output with results saved
waza run eval.yaml -c ./fixtures -v -o results.json

# Filter to specific tasks
waza run eval.yaml -t "task-name-1" -t "task-name-2"

# Parallel execution
waza run eval.yaml --parallel --workers 4

# Save per-task transcripts
waza run eval.yaml --transcript-dir ./transcripts
waza init

Initialize a new evaluation suite with a compliant directory structure.

bash
# Initialize in current directory
waza init

# Initialize in a named directory
waza init my-eval-suite

Creates: eval.yaml, tasks/ with example task, fixtures/ with example fixture.

waza generate

Generate an eval suite from an existing SKILL.md file.

bash
# Generate eval from SKILL.md
waza generate path/to/SKILL.md

# Specify output directory
waza generate SKILL.md --output-dir ./my-eval

Parses YAML frontmatter (name, description) and creates eval.yaml, starter tasks, and fixtures.

waza compare

Compare results from multiple evaluation runs side by side.

bash
# Compare two result files
waza compare run1.json run2.json

# Compare three or more
waza compare gpt4.json claude.json gemini.json

# JSON output
waza compare run1.json run2.json --format json

Shows per-task score deltas, pass rate differences, and aggregate statistics.

waza dev

Iteratively improve SKILL.md frontmatter compliance with automated scoring.

bash
# Score current skill and suggest improvements
waza dev skills/my-skill

# Target high compliance level
waza dev skills/my-skill --target high

# Auto-apply improvements without prompts
waza dev skills/my-skill --target medium --auto --max-iterations 3

Compliance Levels:

  • Low (< 150 chars or no triggers) — Minimal description
  • Medium (150+ chars, has triggers) — Basic trigger coverage
  • Medium-High (+ anti-triggers) — Routing clarity improved
  • High (+ routing markers like INVOKES/FOR SINGLE OPERATIONS) — Full compliance

Scoring Checks:

  • Description length (150+ chars required, 1024 max)
  • Trigger phrases (USE FOR: patterns)
  • Anti-trigger phrases (DO NOT USE FOR: patterns)
  • Routing clarity markers (WORKFLOW SKILL, INVOKES:, etc.)
  • Token budget (500 soft limit, 5000 hard limit)

Coming Soon: Trigger accuracy tests (#36), --skip-integration (#37), --fast (#38), improvement suggestions engine (#34).

Evaluation Spec Format

yaml
name: my-eval
skill: my-skill
version: "1.0"
executor: mock          # or copilot-sdk
tasks:
  - id: task-1
    name: "Describe the task"
    prompt: "Your prompt to the agent"
    expected: "Expected behavior"
    validators:
      - type: code
        config:
          language: go
      - type: text
        config:
          pattern: "expected pattern"

Engines

EngineUseDescription
mockTestingReturns canned responses for validator development
copilot-sdkProductionExecutes via Copilot CLI SDK

Validators

ValidatorWhat it checks
codeCode compiles / passes syntax check
regexOutput matches regex pattern

Configuration

SettingFlagDefault
Fixtures dir--context-dir./fixtures
Output file--output(none)
Verbose--verbosefalse
Parallel--parallelfalse
Workers--workersCPU count
Transcript dir--transcript-dir(none)

Scoring Quick Reference

Each task produces an EvaluationOutcome with:

FieldDescription
score0.0–1.0 normalized score
passBoolean pass/fail
validator_resultsPer-validator details
durationExecution time

© microsoft, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/waza of microsoft/waza.

Open the folder on GitHubat commit 774df00

Compare with similar skills

Waza Skill Evaluator next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Waza Skill Evaluator compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Waza Skill Evaluator this skillmicrosoft/waza1.4k—~2kAutomated safety check: PassMIT
Skill CreatorZS520L/HanakoPro102—~7.4kAutomated safety check: PassApache-2.0
Skill Creatorhimself65/finance-skills3.4k—~3.8kAutomated safety check: PassMIT
Skill Eval ImproveArenukvern/mcp_flutter386—~2.4kAutomated safety check: PassMIT
Skill CreatorAzure/azqr79589 repos~8.2kAutomated safety check: PassApache-2.0
Darwin Skill Optimizeralchaincyf/darwin-skill6.2k1 repos~4.7kAutomated safety check: PassMIT

Similar skills

  • Skill Creator

    ZS520L/HanakoPro

    Create new skills, modify and improve existing skills, and measure skill performance.

    102 GitHub stars~7.4k tokensUpdated 4 mo ago
    Agent WorkflowsAuto-check passed
  • Skill Creator

    himself65/finance-skills

    Create, improve, and evaluate agent skills (SKILL.md plus reference files).

    3.4k GitHub stars~3.8k tokensUpdated 3 days ago
    Agent WorkflowsAuto-check passed
  • Skill Eval Improve

    Arenukvern/mcp_flutter

    Improves Agent Skills via validate → rule-based eval cases → plugin-eval → prompt evals → bounded edits with held-out gates.

    386 GitHub stars~2.4k tokensUpdated 5 days ago
    Agent WorkflowsAuto-check passed
  • Skill Creator

    Azure/azqr

    Official

    Create new skills, modify and improve existing skills, and measure skill performance.

    795 GitHub starsUsed in 89 repos~8.2k tokens
    Agent WorkflowsAuto-check passed
  • Darwin Skill Optimizer

    alchaincyf/darwin-skill

    Scores SKILL.md files on a nine-dimension rubric, then improves them in a keep-or-revert loop with independent judge agents, test prompts, git history and human checkpoints.

    6.2k GitHub starsUsed in 1 repo~4.7k tokens
    Agent WorkflowsAuto-check passed
  • Skill Release Gate

    rohitg00/ai-engineering-from-scratch

    Evaluates an Agent Skill bundle before release for structure, trigger quality, artifact improvement, script correctness, safety, installed-tree integrity and host portability.

    66k GitHub stars~1k tokensUpdated 2 days ago
    Agent WorkflowsAuto-check passed

More from microsoft/waza

All 16 skills in this repo
  • Squad Commands Menu

    microsoft/waza

    Official

    Shows a categorized, interactive menu of common Squad operations, such as install, upgrade and team management, and collects arguments before running anything.

    1.4k GitHub starsUsed in 1 repo~2.7k tokens
    Auto-check passed
  • Official

    Shared collaboration rules for a team of squad agents covering worktree awareness, writing decisions to an inbox, cross-agent requests and reviewer lockout.

    1.4k GitHub starsUsed in 4 repos~500 tokens
    Auto-check passed
  • Official

    Walks through releasing a new version of the waza azd extension: changelog from commits, semver bump with your confirmation, and a release PR.

    1.4k GitHub stars~1.5k tokensUpdated 2 days ago
    Auto-check passed
  • Official

    Dev-first branching model for the Squad project: feature work branches from dev, issue branches follow a naming rule and parallel issues use git worktrees.

    1.4k GitHub starsUsed in 4 repos~1.5k tokens
    Auto-check passed
  • Reviewer Protocol

    microsoft/waza

    Official

    Reviewer rejection workflow and strict lockout semantics. An agent skill from microsoft/waza.

    1.4k GitHub starsUsed in 4 repos~1.1k tokens
    Auto-check passed
  • Waza Interactive

    microsoft/waza

    Official

    Walks you through creating, running and reading waza evals for an agent skill, then proposes concrete fixes when tasks fail or the score is low.

    1.4k GitHub stars~1.3k tokensUpdated 2 days ago
    Auto-check passed

Questions about Waza Skill Evaluator

What does Waza Skill Evaluator do?

Evaluates agent skills with a Go CLI that runs YAML-defined benchmarks, compares runs and scores the quality of SKILL.md frontmatter. waza is a Go command-line tool for testing agent skills. You describe test cases in a YAML eval spec, run them with waza run against an agent engine (the mock engine by default, with fixtures supplied through --context-dir), and score the results with pluggable validators, which the description lists as code and regex checks.

When should I use Waza Skill Evaluator?

Waza Skill Evaluator fits situations like: benchmarking an agent skill against a set of YAML test cases; generating a starter eval suite from an existing SKILL.md; comparing two evaluation runs to see which tasks improved; scoring a skill's frontmatter and description for trigger coverage.

How do I install Waza Skill Evaluator in Claude Code?

Run `npx skills add microsoft/waza --skill waza -a claude-code`. Or copy the skill folder (skills/waza in microsoft/waza) into .claude/skills/waza in your project. Claude Code loads it when a task matches its description.

How do I install Waza Skill Evaluator in Codex?

Run `npx skills add microsoft/waza --skill waza -a codex`. Or copy the skill folder (skills/waza in microsoft/waza) into .agents/skills/waza in your project. Codex loads it when a task matches its description.

Can I use Waza Skill Evaluator in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add microsoft/waza --skill waza -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/waza, .gemini/skills/waza, .github/skills/waza and .opencode/skills/waza in your project.

What does Waza Skill Evaluator need to run?

SKILL.md names no scripts, command-line tools or credentials: Waza Skill Evaluator is instructions for the agent only. Our summary lists: The waza command-line tool; An agent engine such as the Copilot SDK executor, or the built-in mock engine for dry runs.

Does Waza Skill Evaluator access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Waza Skill Evaluator safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Waza Skill Evaluator use?

Waza Skill Evaluator is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Waza Skill Evaluator use?

About 2k tokens (SKILL.md is roughly 7.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Waza Skill Evaluator?

Skills that share tags, products or a category with Waza Skill Evaluator: Skill Creator (ZS520L/HanakoPro, 102 stars), Skill Creator (himself65/finance-skills, 3.4k stars), Skill Eval Improve (Arenukvern/mcp_flutter, 386 stars) and Skill Creator (Azure/azqr, 795 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Waza Skill Evaluator?

microsoft (a GitHub organization, an official publisher) maintains it in microsoft/waza, which has 1,403 GitHub stars. The repository holds 16 skills in this directory. The repository was last updated on October 6, 2026.

Source: microsoft/waza on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.