Official agent skill

Waza Interactive

by microsoft in microsoft/waza

Walks you through creating, running and reading waza evals for an agent skill, then proposes concrete fixes when tasks fail or the score is low.

OfficialMITAuto-check passedAgent Workflows

Install Waza Interactive

skills CLI
$ npx skills add microsoft/waza --skill waza-interactive -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install microsoft/waza waza-interactive --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/microsoft/waza.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/waza-interactive .claude/skills/waza-interactive && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
waza-interactive
GitHub stars
1.4k
Token cost
~1.3k tokens
SKILL.md length
636 words
Files
2
Skills in repo
16
Repo updated
First seen
Licence
MIT

At a glance

Walks you through creating, running and reading waza evals for an agent skill, then proposes concrete fixes when tasks fail or the score is low.

  • Works in 8 steps: Ask which skill to evaluate — get the… → Call waza_eval_list to check for… → If none exist, run waza init via… → …
  • Setting up an eval suite for a skill you are building
  • SKILL.md covers Available MCP Tools, Scenario 1: Create a New Eval, Scenario 2: Run and Interpret… and Scenario 3: Compare Models, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

The agent acts as a conversational partner for waza, a framework for measuring how well agent skills perform. It works through waza's MCP tools, which list eval suites, fetch an eval spec, validate the YAML, start a benchmark run, poll or cancel it, summarize scores, show per-task results and check a skill for compliance.

For a new eval it asks which skill to test, checks for existing suites, runs `waza init` in the terminal if none exist, explains the generated `eval.yaml` and helps write three to five tasks covering the happy path, an edge case and error handling. After a run it reports pass rate, weighted score and duration, treats 90% and above as strong and anything under 70% as a serious problem, and digs into failed tasks when the pass rate drops below 80%. The excerpt also begins a model-comparison scenario but is cut off there.

When your agent uses it

  • Setting up an eval suite for a skill you are building
  • Running waza benchmarks and making sense of the pass rate and weighted score
  • Finding out why particular eval tasks keep failing
  • Comparing how two models score on the same skill eval

Example prompts

  • “Create an eval suite for my pdf-extractor skill in ./skills/pdf-extractor.”
  • “Run the evals for the changelog skill and explain why the pass rate is under 80%.”
  • “Compare gpt-4o and claude-sonnet-4 on the code-review eval.”
  • “Is my release-notes skill ready to ship?”

Requirements

  • The waza CLI
  • Access to the waza MCP tools

Workflow steps

8 steps, taken from the first numbered list in SKILL.md.

  1. Ask which skill to evaluate — get the skill name and path
  2. Call waza_eval_list to check for existing evals for this skill
  3. If none exist, run waza init via terminal to scaffold
  4. Explain the generated eval.yaml structure — name, skill, executor, tasks
  5. Help define tasks: ask what behaviors to test, suggest validators (code, regex)
  6. For each task, help write the prompt and expected output
  7. Call waza_eval_validate to confirm the YAML is valid
  8. Suggest running with waza_eval_run to verify the first task passes

What it can do on your machine

Read from SKILL.md and the folder at commit 774df00. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Waza Interactive loads about 1.3k tokens when it runs. Until then it costs about 102 tokens; SKILL.md has 636 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~102
When it runs · the whole SKILL.md, loaded when a task matches
~1.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from microsoft/waza at commit 774df00, republished under its MIT licence (© microsoft). 636 words, ~1,323 tokens.

Download SKILL.mdSave it as .claude/skills/waza-interactive/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
waza-interactive
description
Interactive workflow partner for creating, testing, and improving AI agent skills with waza. USE FOR: run my evals, check my skill, compare models, create eval suite, debug failing tests, is my skill ready, ship readiness, interpret results, improve score. DO NOT USE FOR: general coding, non-skill work, writing skill content (use skill-authoring), improving frontmatter only (use sensei).

Waza Interactive

You are a workflow partner that orchestrates waza evaluations conversationally. Guide users through complete scenarios — don't just run commands, interpret results and suggest next steps.

Available MCP Tools

Call these tools to execute waza operations:

ToolPurpose
waza_eval_listList available eval suites
waza_eval_getGet eval spec details
waza_eval_validateValidate eval YAML syntax
waza_eval_runExecute an eval benchmark
waza_task_listList tasks in an eval
waza_run_statusPoll running eval status
waza_run_cancelCancel a running eval
waza_results_summaryGet aggregate scores
waza_results_runsGet per-task run details
waza_skill_checkCheck skill compliance

Scenario 1: Create a New Eval

When user wants to create an eval suite for their skill:

  1. Ask which skill to evaluate — get the skill name and path
  2. Call waza_eval_list to check for existing evals for this skill
  3. If none exist, run waza init <directory> via terminal to scaffold
  4. Explain the generated eval.yaml structure — name, skill, executor, tasks
  5. Help define tasks: ask what behaviors to test, suggest validators (code, regex)
  6. For each task, help write the prompt and expected output
  7. Call waza_eval_validate to confirm the YAML is valid
  8. Suggest running with waza_eval_run to verify the first task passes

Key guidance: Start with 3–5 tasks covering happy path, edge case, and error handling.

Scenario 2: Run and Interpret Results

When user wants to run evals and understand scores:

  1. Call waza_eval_run with the eval spec path and context dir
  2. Poll waza_run_status until complete (check every 10s)
  3. Call waza_results_summary to get aggregate scores
  4. Interpret the results for the user:
    • Pass rate — percentage of tasks that passed all validators
    • Weighted score — 0.0–1.0 aggregate across all tasks
    • Duration — total and per-task execution time
  5. If pass rate < 80%, identify which tasks failed and why
  6. Call waza_results_runs for per-task details on failures
  7. Suggest specific improvements: prompt rewording, validator tuning, fixture updates

Thresholds: ≥90% pass rate = strong, 70–89% = needs work, <70% = significant issues.

Scenario 3: Compare Models

When user wants to compare model performance:

  1. Ask which models to compare (e.g., gpt-4o vs claude-sonnet-4)
  2. Call waza_eval_run with model A — save results
  3. Call waza_eval_run with model B — save results
  4. Compare results side by side:
    • Per-task pass/fail differences
    • Score deltas (which model scores higher on which tasks)
    • Duration differences (speed vs quality tradeoff)
  5. Provide a recommendation: which model is better for this skill and why
  6. Suggest next steps: try a third model, tune prompts for the weaker model, or adjust validators

Guidance: Run each model 2–3 times to account for variance before drawing conclusions.

Show full SKILL.md (214 more words)Show less

Scenario 4: Debug a Failing Skill

When user's skill is failing evals or behaving unexpectedly:

  1. Call waza_skill_check to verify skill compliance (frontmatter, triggers, token count)
  2. If compliance issues found, fix those first — they affect routing
  3. Call waza_eval_run with --verbose and --transcript-dir flags
  4. Call waza_results_runs to get per-task failure details
  5. Analyze failure patterns:
    • All tasks fail → prompt or fixture issue, check skill instructions
    • Some tasks fail → specific edge cases, review failed task prompts
    • Validator failures → regex too strict, code validator language mismatch
  6. Suggest targeted fixes based on the pattern
  7. Re-run with waza_eval_run to verify the fix

Scenario 5: Ship Readiness Check

When user asks "is my skill ready?" or wants a pre-ship checklist:

  1. Call waza_skill_check — verify compliance score ≥ medium-high
  2. Call waza_eval_validate — confirm eval YAML is valid
  3. Call waza_eval_run — execute full eval suite
  4. Call waza_results_summary — check aggregate scores
  5. Render the readiness verdict:
SHIP READINESS CHECKLIST:
☐ Skill compliance: [score] (need: medium-high+)
☐ Eval YAML valid: [yes/no]
☐ Pass rate: [X]% (need: ≥90%)
☐ Weighted score: [X.XX] (need: ≥0.85)
☐ No task timeouts
☐ Consistent across 2+ runs

VERDICT: [READY / NOT READY — fix items marked ✗]
  1. If NOT READY, route to the appropriate scenario (Scenario 4 for failures, Scenario 1 for missing evals)

Conversation Style

  • Always explain why before what — context before commands
  • After every tool call, interpret the result in plain language
  • When something fails, diagnose before suggesting fixes
  • Offer the next logical step — don't wait to be asked
  • Use the checklist format for multi-step validations

© microsoft, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in skills/waza-interactive of microsoft/waza.

  • SKILL.md
  • tests/eval.yaml

Open the folder on GitHubat commit 774df00

Compare with similar skills

Waza Interactive next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Waza Interactive compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Waza Interactive this skillmicrosoft/waza1.4k—~1.3kAutomated safety check: PassMIT
Skill ForgeAgriciDaniel/skill-forge177—~1.9kAutomated safety check: NotesMIT
Skill CreatorZS520L/HanakoPro102—~7.4kAutomated safety check: PassApache-2.0
Octocode Graph Eval Loopbgauryy/octocode946—~1.6kAutomated safety check: PassMIT
Create Skilldavekilleen/Dex493—~2.9kAutomated safety check: PassCustom licence
Skill Creatorhimself65/finance-skills3.4k—~3.8kAutomated safety check: PassMIT

Similar skills

  • Skill Forge

    AgriciDaniel/skill-forge

    Ultimate Claude Code skill creator and architect. An agent skill from AgriciDaniel/skill-forge.

    177 GitHub stars~1.9k tokensUpdated 5 mo ago
    Agent WorkflowsAuto-check: notes
  • Skill Creator

    ZS520L/HanakoPro

    Create new skills, modify and improve existing skills, and measure skill performance.

    102 GitHub stars~7.4k tokensUpdated 4 mo ago
    Agent WorkflowsAuto-check passed
  • Octocode Graph Eval Loop

    bgauryy/octocode

    Runs a measurable keep-or-discard improvement loop against a runnable sensor, from framing a goal and KPI through baseline, judging and held-out verification.

    946 GitHub stars~1.6k tokensUpdated 3 days ago
    Agent WorkflowsAuto-check passed
  • Create Skill

    davekilleen/Dex

    Author a new Dex skill that actually fires and passes the quality bar.

    493 GitHub stars~2.9k tokensUpdated 5 days ago
    Agent WorkflowsAuto-check passed
  • Skill Creator

    himself65/finance-skills

    Create, improve, and evaluate agent skills (SKILL.md plus reference files).

    3.4k GitHub stars~3.8k tokensUpdated 2 days ago
    Agent WorkflowsAuto-check passed
  • Benchmark Agents

    vercel/vercel-plugin

    Official

    Advanced AI agent benchmark scenarios that push Vercel's cutting-edge platform features — Workflow SDK, AI Gateway, MCP, Chat SDK, Queues, Flags, Sandbox, and multi-agent orchestration.

    301 GitHub stars~3.6k tokensUpdated today
    Agent WorkflowsAuto-check passed

More from microsoft/waza

All 16 skills in this repo
  • Squad Commands Menu

    microsoft/waza

    Official

    Shows a categorized, interactive menu of common Squad operations, such as install, upgrade and team management, and collects arguments before running anything.

    1.4k GitHub starsUsed in 1 repo~2.7k tokens
    Auto-check passed
  • Official

    Shared collaboration rules for a team of squad agents covering worktree awareness, writing decisions to an inbox, cross-agent requests and reviewer lockout.

    1.4k GitHub starsUsed in 4 repos~500 tokens
    Auto-check passed
  • Official

    Walks through releasing a new version of the waza azd extension: changelog from commits, semver bump with your confirmation, and a release PR.

    1.4k GitHub stars~1.5k tokensUpdated yesterday
    Auto-check passed
  • Official

    Dev-first branching model for the Squad project: feature work branches from dev, issue branches follow a naming rule and parallel issues use git worktrees.

    1.4k GitHub starsUsed in 4 repos~1.5k tokens
    Auto-check passed
  • Waza Skill Evaluator

    microsoft/waza

    Official

    Evaluates agent skills with a Go CLI that runs YAML-defined benchmarks, compares runs and scores the quality of SKILL.md frontmatter.

    1.4k GitHub stars~2k tokensUpdated yesterday
    Auto-check passed
  • Reviewer Protocol

    microsoft/waza

    Official

    Reviewer rejection workflow and strict lockout semantics. An agent skill from microsoft/waza.

    1.4k GitHub starsUsed in 4 repos~1.1k tokens
    Auto-check passed

Questions about Waza Interactive

What does Waza Interactive do?

Walks you through creating, running and reading waza evals for an agent skill, then proposes concrete fixes when tasks fail or the score is low. The agent acts as a conversational partner for waza, a framework for measuring how well agent skills perform. It works through waza's MCP tools, which list eval suites, fetch an eval spec, validate the YAML, start a benchmark run, poll or cancel it, summarize scores, show per-task results and check a skill for compliance.

When should I use Waza Interactive?

Waza Interactive fits situations like: setting up an eval suite for a skill you are building; running waza benchmarks and making sense of the pass rate and weighted score; finding out why particular eval tasks keep failing; comparing how two models score on the same skill eval.

How do I install Waza Interactive in Claude Code?

Run `npx skills add microsoft/waza --skill waza-interactive -a claude-code`. Or copy the skill folder (skills/waza-interactive in microsoft/waza) into .claude/skills/waza-interactive in your project. Claude Code loads it when a task matches its description.

How do I install Waza Interactive in Codex?

Run `npx skills add microsoft/waza --skill waza-interactive -a codex`. Or copy the skill folder (skills/waza-interactive in microsoft/waza) into .agents/skills/waza-interactive in your project. Codex loads it when a task matches its description.

Can I use Waza Interactive in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add microsoft/waza --skill waza-interactive -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/waza-interactive, .gemini/skills/waza-interactive, .github/skills/waza-interactive and .opencode/skills/waza-interactive in your project.

What does Waza Interactive need to run?

SKILL.md names no scripts, command-line tools or credentials: Waza Interactive is instructions for the agent only. Our summary lists: The waza CLI; Access to the waza MCP tools.

Does Waza Interactive access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Waza Interactive safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Waza Interactive use?

Waza Interactive is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Waza Interactive use?

About 1.3k tokens (SKILL.md is roughly 5.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Waza Interactive?

Skills that share tags, products or a category with Waza Interactive: Skill Forge (AgriciDaniel/skill-forge, 177 stars), Skill Creator (ZS520L/HanakoPro, 102 stars), Octocode Graph Eval Loop (bgauryy/octocode, 946 stars) and Create Skill (davekilleen/Dex, 493 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Waza Interactive?

microsoft (a GitHub organization, an official publisher) maintains it in microsoft/waza, which has 1,400 GitHub stars. The repository holds 16 skills in this directory. The repository was last updated on October 6, 2026.

Source: microsoft/waza on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.