Agent skill

Waxa Eval

by mizchi in mizchi/skills

A skill your agent uses when iterating on a skill's prompt with the waxa CLI (https://github.com/mizchi/skills/tree/main/tools/waxa) — authoring scenarios, choosing graders, interpreting…

No licenceAuto-check passed

Install Waxa Eval

skills CLI
$ npx skills add mizchi/skills --skill waxa-eval -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install mizchi/skills waxa-eval --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/mizchi/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/waxa-eval .claude/skills/waxa-eval && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
waxa-eval
GitHub stars
359
Token cost
~4.1k tokens
SKILL.md length
1,769 words
Files
2
Skills in repo
69
Repo updated
First seen
Licence
None found

At a glance

A skill your agent uses when iterating on a skill's prompt with the waxa CLI (https://github.com/mizchi/skills/tree/main/tools/waxa) — authoring scenarios, choosing graders, interpreting…

  • Works in 2 steps: A surface text grader with deliberately… → An llm grader with a multi-clause…
  • Iterating on a skills prompt with the waxa CLI (https://github.com/mizchi/skills/tree/main/tools/waxa) — authoring scenarios
  • SKILL.md covers When to invoke, The four-stage iteration pattern, Authoring scenarios… and Grader selection, plus 8 more sections
  • Calls npx and claude

What it does

Waxa Eval is an agent skill from mizchi/skills. Use when iterating on a skill's prompt with the waxa CLI (https://github.com/mizchi/skills/tree/main/tools/waxa) — authoring scenarios, choosing graders, interpreting unclear-points, advancing a ledger, and judging convergence. Encodes the four-stage iteration pattern observed in real iter loops (structural fix → grader breadth → surface-form coverage → residual unclear) and the scenario-design pitfalls (blank-slate executor's network limit, prompt expectation explicitness, regex coverage of Japanese/English…

Its SKILL.md is about 4.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `README.md`).

It works with GitHub. The repository describes itself as: Agent skills by mizchi, distributed via APM.

When your agent uses it

  • Iterating on a skills prompt with the waxa CLI (https://github.com/mizchi/skills/tree/main/tools/waxa) — authoring scenarios
  • Choosing graders
  • Interpreting unclear-points
  • Advancing a ledger

Example prompts

  • “/waxa-eval”

Requirements

  • Node.js

Workflow steps

2 steps, taken from the first numbered list in SKILL.md.

  1. A surface text grader with deliberately broad alternation. Goal: deterministic, fast, catches obvious misses.
  2. An llm grader with a multi-clause rubric. Goal: catches semantic equivalents and judges nuance.

What it can do on your machine

Read from SKILL.md and the folder at commit 62f5808. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • npx
    • claude

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com
    • agentskills.io

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Waxa Eval loads about 4.1k tokens when it runs. Until then it costs about 169 tokens; SKILL.md has 1,769 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~169
When it runs · the whole SKILL.md, loaded when a task matches
~4.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

Without a licence we can't republish the file, so here is its outline and opening line. It has 1,769 words (~4,110 tokens).

“Empirical evaluation loop for skill prompts, codified from real iter runs. This skill is the operating manual for waxa; the CLI itself lives in tools/waxa/ (see its README for argument-level reference).”

— opening of SKILL.md by mizchi
name
waxa-eval

Read the full SKILL.md on GitHub

Files

SKILL.md and 1 other file in waxa-eval of mizchi/skills.

  • SKILL.md
  • README.md

Open the folder on GitHubat commit 62f5808

Compare with similar skills

Waxa Eval next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Waxa Eval compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Waxa Eval this skillmizchi/skills359—~4.1kAutomated safety check: PassNone
PR Babysitteropeninterpreter/openinterpreter69k3 repos~4.2kAutomated safety check: PassApache-2.0
Diagnosing Superpowers Sessionsobra/superpowers297k3 repos~1.7kAutomated safety check: PassMIT
Greplooponyx-dot-app/onyx32k4 repos~3.3kAutomated safety check: PassMIT
GitHub Deep Researchbytedance/deer-flow84k4 repos~1.3kAutomated safety check: PassMIT
Update V8 Versionopeninterpreter/openinterpreter69k2 repos~845Automated safety check: PassApache-2.0

Similar skills

  • PR Babysitter

    openinterpreter/openinterpreter

    Watches an open GitHub pull request until it merges, handling review comments, diagnosing CI failures and retrying flaky checks along the way.

    69k GitHub starsUsed in 3 repos~4.2k tokens
    DevelopmentAuto-check passed
  • Investigates a session where Superpowers went wrong, reads the transcripts on disk and produces an evidence-cited report, optionally prepared as a bug report for the maintainers.

    297k GitHub starsUsed in 3 repos~1.7k tokens
    Agent WorkflowsAuto-check passed
  • Greploop

    onyx-dot-app/onyx

    Iteratively improves a PR (GitHub), MR (GitLab), or shelved changelist (Perforce) until Greptile gives it a 5/5 confidence score with zero unresolved comments.

    32k GitHub starsUsed in 4 repos~3.3k tokens
    DevelopmentAuto-check passed
  • GitHub Deep Research

    bytedance/deer-flow

    Researches a GitHub repository over four rounds using the GitHub API and web search, then writes a structured markdown report with timeline, metrics and Mermaid diagrams.

    84k GitHub starsUsed in 4 repos~1.3k tokens
    Research & ScienceAuto-check passed
  • Update V8 Version

    openinterpreter/openinterpreter

    Bumps the pinned v8 and rusty_v8 versions in Codex, validates the release-candidate path with the v8-canary check, and traces failures to upstream build changes.

    69k GitHub starsUsed in 2 repos~845 tokens
    DevOps & CloudAuto-check passed
  • Last30days

    mvanhorn/last30days-skill

    Research what people actually say about any topic in the last 30 days.

    64k GitHub stars~7.9k tokensUpdated today
    Research & ScienceAuto-check: notes

More from mizchi/skills

All 69 skills in this repo
  • Publish Agent Skill

    mizchi/skills

    Package, publish, verify and update an agent skill through a Claude Code plugin marketplace, APM or the npx skills CLI.

    359 GitHub stars~1.7k tokensUpdated 7 days ago
    Auto-check passed
  • AI Index

    mizchi/skills

    Method and tooling for measuring how AI-generated a piece of prose reads, in Japanese or English.

    359 GitHub stars~3.6k tokensUpdated 7 days ago
    Auto-check passed
  • Cloudflare Deploy

    mizchi/skills

    Deploy applications and infrastructure to Cloudflare with the cf CLI and typed cloudflare.config.ts.

    359 GitHub stars~2.9k tokensUpdated 7 days ago
    Auto-check: notes
  • Review Image

    mizchi/skills

    Review screenshots or other images with OpenRouter vision models via bundled Deno scripts.

    359 GitHub stars~1.3k tokensUpdated 7 days ago
    Auto-check passed
  • Post-generation safety checks for sqlc-gen-moonbit + Cloudflare D1.

    359 GitHub stars~995 tokensUpdated 7 days ago
    Auto-check passed
  • Apm Usage

    mizchi/skills

    Reference for APM (Agent Package Manager) — apm.yml syntax, install / uninstall / update commands, target detection, lockfile workflow.

    359 GitHub stars~1.9k tokensUpdated 7 days ago
    Auto-check passed

Works with

Questions about Waxa Eval

What does Waxa Eval do?

A skill your agent uses when iterating on a skill's prompt with the waxa CLI (https://github.com/mizchi/skills/tree/main/tools/waxa) — authoring scenarios, choosing graders, interpreting…. Waxa Eval is an agent skill from mizchi/skills.com/mizchi/skills/tree/main/tools/waxa) — authoring scenarios, choosing graders, interpreting unclear-points, advancing a ledger, and judging convergence.

When should I use Waxa Eval?

Waxa Eval fits situations like: iterating on a skills prompt with the waxa CLI (https://github.com/mizchi/skills/tree/main/tools/waxa) — authoring scenarios; choosing graders; interpreting unclear-points; advancing a ledger.

How do I install Waxa Eval in Claude Code?

Run `npx skills add mizchi/skills --skill waxa-eval -a claude-code`. Or copy the skill folder (waxa-eval in mizchi/skills) into .claude/skills/waxa-eval in your project. Claude Code loads it when a task matches its description.

How do I install Waxa Eval in Codex?

Run `npx skills add mizchi/skills --skill waxa-eval -a codex`. Or copy the skill folder (waxa-eval in mizchi/skills) into .agents/skills/waxa-eval in your project. Codex loads it when a task matches its description.

Can I use Waxa Eval in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mizchi/skills --skill waxa-eval -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/waxa-eval, .gemini/skills/waxa-eval, .github/skills/waxa-eval and .opencode/skills/waxa-eval in your project.

What does Waxa Eval need to run?

Going by SKILL.md and its folder, Waxa Eval needs the command-line tools its instructions call (npx and claude). Our summary lists: Node.js.

Does Waxa Eval access the network?

SKILL.md names 2 domains. As links in the text: github.com and agentskills.io. This is read from the text; nothing was executed.

Is Waxa Eval safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Waxa Eval use?

No licence was found for Waxa Eval or its repository. Without one, default copyright applies: ask the author before reusing or redistributing it.

How many tokens does Waxa Eval use?

About 4.1k tokens (SKILL.md is roughly 16k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Waxa Eval?

Skills that share tags, products or a category with Waxa Eval: PR Babysitter (openinterpreter/openinterpreter, 69k stars), Diagnosing Superpowers Sessions (obra/superpowers, 297k stars), Greploop (onyx-dot-app/onyx, 32k stars) and GitHub Deep Research (bytedance/deer-flow, 84k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Waxa Eval?

mizchi (a GitHub user) maintains it in mizchi/skills, which has 359 GitHub stars. The repository holds 69 skills in this directory. The repository was last updated on October 2, 2026.

Source: mizchi/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.