Agent skill

Windmill AI Evals

by windmill-labs in windmill-labs/windmill

Writes and runs black-box benchmark cases for Windmill's flow, app, script, CLI and global AI generation modes, including before-and-after comparisons.

Custom licenceAuto-check: notesAI & LLM Engineering

Install Windmill AI Evals

skills CLI
$ npx skills add windmill-labs/windmill --skill ai-evals -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install windmill-labs/windmill ai-evals --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/windmill-labs/windmill.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/ai-evals .claude/skills/ai-evals && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ai-evals
GitHub stars
18k
Token cost
~969 tokens
SKILL.md length
426 words
Files
1
Skills in repo
42
Repo updated
First seen
Licence
Custom licence

At a glance

Writes and runs black-box benchmark cases for Windmill's flow, app, script, CLI and global AI generation modes, including before-and-after comparisons.

  • Works in 5 steps: Write prompts like a real user request. → Prefer behavior, inputs, constraints,… → Keep deterministic validation narrow and… → …
  • Adding a benchmark case for a Windmill AI generation mode
  • SKILL.md covers Running benchmarks and Authoring core rules
  • Calls bun

What it does

The ai_evals directory holds a benchmark runner that always tests the current production prompts, tools and guidance in your checkout. Each attempt goes through the real production path, then deterministic validation, then LLM judging. The aim is realistic user requests, not pinning one exact implementation. You run it with bun: install once, list model aliases with the CLI, then run chosen cases against chosen models.

The frontend modes send model calls through a Windmill backend's AI proxy, so a reachable backend is needed, set through environment variables for its URL and workspace. Reuse an existing workspace, because community builds cap how many exist, and the only side effect is an upserted AI resource. Provider keys come from ai_evals/.env, and the judge is a separate Anthropic call whatever model is under test.

Authoring rules: write prompts like a real user request, prefer behavior, inputs, constraints and outcomes over internals, keep deterministic validation narrow and hard, put semantic expectations in judgeChecklist, and use expected fixtures only when exact structure matters. The excerpt is cut off after the prompt-writing examples.

When your agent uses it

  • Adding a benchmark case for a Windmill AI generation mode
  • Running before-and-after benchmarks for a copilot or AI chat change
  • Editing the judge checklist or fixtures of an existing eval case
  • Rewriting an eval prompt so it reads like a real user request

Example prompts

  • “Add an eval case for flow mode: route support requests based on customer tier.”
  • “Run the global benchmark cases before and after my copilot change and compare the results.”
  • “Rewrite this eval prompt so it does not name any Windmill internals.”

Requirements

  • Bun and a Windmill checkout with the ai_evals directory
  • A reachable Windmill backend and an existing workspace
  • Model provider keys in ai_evals/.env

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Write prompts like a real user request.
  2. Prefer behavior, inputs, constraints, and outcomes over internal implementation.
  3. Keep deterministic validation narrow and hard.
  4. Put semantic expectations in judgeChecklist.
  5. Use expected fixtures only when exact structure really matters.

What it can do on your machine

Read from SKILL.md and the folder at commit fb22e5c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • bun

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Windmill AI Evals loads about 969 tokens when it runs. Until then it costs about 60 tokens; SKILL.md has 426 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~60
When it runs · the whole SKILL.md, loaded when a task matches
~969

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:39
    - Provider keys live in `ai_evals/.env` and are auto-loaded by bun. The judge is a

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

Its licence (Custom licence) doesn't allow us to republish the file, so here is its outline and opening line. It has 426 words (~969 tokens).

“ai_evals/ is a black-box benchmark runner for the Windmill AI generation modes: flow, app, script, cli, global. It always tests the current production prompts, tools, and guidance in this checkout. Each attempt runs the real production path, deterministic validation, then…”

— opening of SKILL.md by windmill-labs, Custom licence
name
ai-evals

Read the full SKILL.md on GitHub

Files

Just SKILL.md in .agents/skills/ai-evals of windmill-labs/windmill.

Open the folder on GitHubat commit fb22e5c

Compare with similar skills

Windmill AI Evals next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Windmill AI Evals compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Windmill AI Evals this skillwindmill-labs/windmill18k—~969Automated safety check: NotesCustom licence
Agent Eval Engineeringlangchain-ai/langchain-skills1.3k—~4kAutomated safety check: PassMIT
Octocode Benchmark Runnerbgauryy/octocode949—~2.1kAutomated safety check: PassMIT
SWE Benchmark Task Adderory/lumen307—~497Automated safety check: PassCustom licence
GAIA Agent Benchmarkingamd/gaia1.6k—~1.8kAutomated safety check: PassMIT
Autocontextgreyhaven-ai/autocontext1.3k—~892Automated safety check: PassApache-2.0

Similar skills

  • Agent Eval Engineering

    langchain-ai/langchain-skills

    Official

    Builds agent evaluations in stages: inspect the repository and traces, agree a Task Spec with you, then build, audit and run a Harbor task with an independent verifier.

    1.3k GitHub stars~4k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Runs blind pairwise comparisons of Octocode against a gh-based baseline over markdown research questions, scored by total characters through the model rather than self-report.

    949 GitHub stars~2.1k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Adds a new task to the bench-swe pipeline from a real GitHub bug-fix issue or pull request, then checks the generated task file and patch.

    307 GitHub stars~497 tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Benchmarks AMD's GAIA agent against Claude Code and across models on quality, honesty, steps, tokens, time and real cost, using gaia eval tasks.

    1.6k GitHub stars~1.8k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Autocontext

    greyhaven-ai/autocontext

    Runs LLM-based rubric judging on agent output and loops revise-and-rejudge rounds until a quality threshold is met.

    1.3k GitHub stars~892 tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Benchflow

    benchflow-ai/benchflow

    Run agent benchmarks, create tasks, analyze results, and manage agents using BenchFlow.

    355 GitHub stars~1.9k tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check: notes

More from windmill-labs/windmill

All 42 skills in this repo
  • Windmill Trigger Type Checklist

    windmill-labs/windmill

    Checklist of every backend, frontend, CLI and capture change needed to add a new TriggerCrud-based trigger type, such as Azure, GCP or Kafka, to Windmill.

    18k GitHub stars~4.7k tokensUpdated today
    Auto-check passed
  • Domain Modeling and Glossary

    windmill-labs/windmill

    Actively challenges vague or conflicting terminology as you design, and keeps a living domain glossary file up to date in real time.

    18k GitHub stars~622 tokensUpdated today
    Auto-check passed
  • Local PR Review

    windmill-labs/windmill

    Runs the same code review locally that GitHub's auto-review actions run on a PR, delegating to a fresh-context subagent so the review isn't biased by the main session's own reasoning.

    18k GitHub stars~995 tokensUpdated today
    Auto-check passed
  • Draft PR with CI Rounds

    windmill-labs/windmill

    Opens a draft GitHub pull request with a conventional title and explicit body, then drives CI review rounds before marking it ready.

    18k GitHub stars~3.4k tokensUpdated today
    Auto-check passed
  • Windmill Dev Page Preview

    windmill-labs/windmill

    Opens the Windmill dev page to preview a flow, script or app, choosing between proxy and direct mode and deciding whether the agent or the runtime starts the wmill dev server.

    18k GitHub stars~1.8k tokensUpdated today
    Auto-check passed
  • Windmill Rust Backend Patterns

    windmill-labs/windmill

    Sets Rust conventions for the Windmill backend: error types, SQLx queries, JSON handling, async rules, module layout and rust-analyzer navigation.

    18k GitHub stars~869 tokensUpdated today
    Auto-check passed

Questions about Windmill AI Evals

What does Windmill AI Evals do?

Writes and runs black-box benchmark cases for Windmill's flow, app, script, CLI and global AI generation modes, including before-and-after comparisons. The ai_evals directory holds a benchmark runner that always tests the current production prompts, tools and guidance in your checkout. Each attempt goes through the real production path, then deterministic validation, then LLM judging.

When should I use Windmill AI Evals?

Windmill AI Evals fits situations like: adding a benchmark case for a Windmill AI generation mode; running before-and-after benchmarks for a copilot or AI chat change; editing the judge checklist or fixtures of an existing eval case; rewriting an eval prompt so it reads like a real user request.

How do I install Windmill AI Evals in Claude Code?

Run `npx skills add windmill-labs/windmill --skill ai-evals -a claude-code`. Or copy the skill folder (.agents/skills/ai-evals in windmill-labs/windmill) into .claude/skills/ai-evals in your project. Claude Code loads it when a task matches its description.

How do I install Windmill AI Evals in Codex?

Run `npx skills add windmill-labs/windmill --skill ai-evals -a codex`. Or copy the skill folder (.agents/skills/ai-evals in windmill-labs/windmill) into .agents/skills/ai-evals in your project. Codex loads it when a task matches its description.

Can I use Windmill AI Evals in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add windmill-labs/windmill --skill ai-evals -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ai-evals, .gemini/skills/ai-evals, .github/skills/ai-evals and .opencode/skills/ai-evals in your project.

What does Windmill AI Evals need to run?

Going by SKILL.md and its folder, Windmill AI Evals needs the command-line tools its instructions call (bun). Our summary lists: Bun and a Windmill checkout with the ai_evals directory; A reachable Windmill backend and an existing workspace; Model provider keys in ai_evals/.env.

Does Windmill AI Evals access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Windmill AI Evals safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Windmill AI Evals use?

Windmill AI Evals has a licence file (the repository's licence) that doesn't match a standard licence. Read it on GitHub before reusing the skill.

How many tokens does Windmill AI Evals use?

About 969 tokens (SKILL.md is roughly 3.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Windmill AI Evals?

Skills that share tags, products or a category with Windmill AI Evals: Agent Eval Engineering (langchain-ai/langchain-skills, 1.3k stars), Octocode Benchmark Runner (bgauryy/octocode, 949 stars), SWE Benchmark Task Adder (ory/lumen, 307 stars) and GAIA Agent Benchmarking (amd/gaia, 1.6k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Windmill AI Evals?

windmill-labs (a GitHub organization) maintains it in windmill-labs/windmill, which has 18,146 GitHub stars. The repository holds 42 skills in this directory. The repository was last updated on October 9, 2026.

Source: windmill-labs/windmill on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.