Agent Eval Engineering
langchain-ai/langchain-skills
Builds agent evaluations in stages: inspect the repository and traces, agree a Task Spec with you, then build, audit and run a Harbor task with an independent verifier.
Writes and runs black-box benchmark cases for Windmill's flow, app, script, CLI and global AI generation modes, including before-and-after comparisons.
$ npx skills add windmill-labs/windmill --skill ai-evals -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install windmill-labs/windmill ai-evals --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/windmill-labs/windmill.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/ai-evals .claude/skills/ai-evals && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "ai-evals" agent skill from https://github.com/windmill-labs/windmill/tree/main/.agents/skills/ai-evals into .claude/skills/ai-evals/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ai-evals", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/windmill-labs/windmill/tree/main/.agents/skills/ai-evalsType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add windmill-labs/windmill --skill ai-evals -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install windmill-labs/windmill ai-evals --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/windmill-labs/windmill.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.agents/skills/ai-evals .agents/skills/ai-evals && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "ai-evals" agent skill from https://github.com/windmill-labs/windmill/tree/main/.agents/skills/ai-evals into .agents/skills/ai-evals/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ai-evals", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add windmill-labs/windmill --skill ai-evals -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install windmill-labs/windmill ai-evals --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/windmill-labs/windmill.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.agents/skills/ai-evals .cursor/skills/ai-evals && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "ai-evals" agent skill from https://github.com/windmill-labs/windmill/tree/main/.agents/skills/ai-evals into .cursor/skills/ai-evals/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ai-evals", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/windmill-labs/windmill.git --path .agents/skills/ai-evals--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add windmill-labs/windmill --skill ai-evals -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install windmill-labs/windmill ai-evals --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/windmill-labs/windmill.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.agents/skills/ai-evals .gemini/skills/ai-evals && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "ai-evals" agent skill from https://github.com/windmill-labs/windmill/tree/main/.agents/skills/ai-evals into .gemini/skills/ai-evals/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ai-evals", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install windmill-labs/windmill ai-evalsInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add windmill-labs/windmill --skill ai-evals -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/windmill-labs/windmill.git skills-src && mkdir -p .github/skills && cp -r skills-src/.agents/skills/ai-evals .github/skills/ai-evals && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "ai-evals" agent skill from https://github.com/windmill-labs/windmill/tree/main/.agents/skills/ai-evals into .github/skills/ai-evals/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ai-evals", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add windmill-labs/windmill --skill ai-evals -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install windmill-labs/windmill ai-evals --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/windmill-labs/windmill.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.agents/skills/ai-evals .opencode/skills/ai-evals && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "ai-evals" agent skill from https://github.com/windmill-labs/windmill/tree/main/.agents/skills/ai-evals into .opencode/skills/ai-evals/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ai-evals", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
ai-evalsWrites and runs black-box benchmark cases for Windmill's flow, app, script, CLI and global AI generation modes, including before-and-after comparisons.
The ai_evals directory holds a benchmark runner that always tests the current production prompts, tools and guidance in your checkout. Each attempt goes through the real production path, then deterministic validation, then LLM judging. The aim is realistic user requests, not pinning one exact implementation. You run it with bun: install once, list model aliases with the CLI, then run chosen cases against chosen models.
The frontend modes send model calls through a Windmill backend's AI proxy, so a reachable backend is needed, set through environment variables for its URL and workspace. Reuse an existing workspace, because community builds cap how many exist, and the only side effect is an upserted AI resource. Provider keys come from ai_evals/.env, and the judge is a separate Anthropic call whatever model is under test.
Authoring rules: write prompts like a real user request, prefer behavior, inputs, constraints and outcomes over internals, keep deterministic validation narrow and hard, put semantic expectations in judgeChecklist, and use expected fixtures only when exact structure matters. The excerpt is cut off after the prompt-writing examples.
5 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit fb22e5c. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
bunFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Windmill AI Evals loads about 969 tokens when it runs. Until then it costs about 60 tokens; SKILL.md has 426 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
- Provider keys live in `ai_evals/.env` and are auto-loaded by bun. The judge is aAutomated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
Its licence (Custom licence) doesn't allow us to republish the file, so here is its outline and opening line. It has 426 words (~969 tokens).
“ai_evals/ is a black-box benchmark runner for the Windmill AI generation modes: flow, app, script, cli, global. It always tests the current production prompts, tools, and guidance in this checkout. Each attempt runs the real production path, deterministic validation, then…”
Just SKILL.md in .agents/skills/ai-evals of windmill-labs/windmill.
Open the folder on GitHubat commit fb22e5c
Windmill AI Evals next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Windmill AI Evals this skillwindmill-labs/windmill | 18k | — | ~969 | Automated safety check: Notes | Custom licence | |
| Agent Eval Engineeringlangchain-ai/langchain-skills | 1.3k | — | ~4k | Automated safety check: Pass | MIT | |
| Octocode Benchmark Runnerbgauryy/octocode | 949 | — | ~2.1k | Automated safety check: Pass | MIT | |
| SWE Benchmark Task Adderory/lumen | 307 | — | ~497 | Automated safety check: Pass | Custom licence | |
| GAIA Agent Benchmarkingamd/gaia | 1.6k | — | ~1.8k | Automated safety check: Pass | MIT | |
| Autocontextgreyhaven-ai/autocontext | 1.3k | — | ~892 | Automated safety check: Pass | Apache-2.0 |
langchain-ai/langchain-skills
Builds agent evaluations in stages: inspect the repository and traces, agree a Task Spec with you, then build, audit and run a Harbor task with an independent verifier.
bgauryy/octocode
Runs blind pairwise comparisons of Octocode against a gh-based baseline over markdown research questions, scored by total characters through the model rather than self-report.
ory/lumen
Adds a new task to the bench-swe pipeline from a real GitHub bug-fix issue or pull request, then checks the generated task file and patch.
amd/gaia
Benchmarks AMD's GAIA agent against Claude Code and across models on quality, honesty, steps, tokens, time and real cost, using gaia eval tasks.
greyhaven-ai/autocontext
Runs LLM-based rubric judging on agent output and loops revise-and-rejudge rounds until a quality threshold is met.
benchflow-ai/benchflow
Run agent benchmarks, create tasks, analyze results, and manage agents using BenchFlow.
windmill-labs/windmill
Checklist of every backend, frontend, CLI and capture change needed to add a new TriggerCrud-based trigger type, such as Azure, GCP or Kafka, to Windmill.
windmill-labs/windmill
Actively challenges vague or conflicting terminology as you design, and keeps a living domain glossary file up to date in real time.
windmill-labs/windmill
Runs the same code review locally that GitHub's auto-review actions run on a PR, delegating to a fresh-context subagent so the review isn't biased by the main session's own reasoning.
windmill-labs/windmill
Opens a draft GitHub pull request with a conventional title and explicit body, then drives CI review rounds before marking it ready.
windmill-labs/windmill
Opens the Windmill dev page to preview a flow, script or app, choosing between proxy and direct mode and deciding whether the agent or the runtime starts the wmill dev server.
windmill-labs/windmill
Sets Rust conventions for the Windmill backend: error types, SQLx queries, JSON handling, async rules, module layout and rust-analyzer navigation.
Categories
Writes and runs black-box benchmark cases for Windmill's flow, app, script, CLI and global AI generation modes, including before-and-after comparisons. The ai_evals directory holds a benchmark runner that always tests the current production prompts, tools and guidance in your checkout. Each attempt goes through the real production path, then deterministic validation, then LLM judging.
Windmill AI Evals fits situations like: adding a benchmark case for a Windmill AI generation mode; running before-and-after benchmarks for a copilot or AI chat change; editing the judge checklist or fixtures of an existing eval case; rewriting an eval prompt so it reads like a real user request.
Run `npx skills add windmill-labs/windmill --skill ai-evals -a claude-code`. Or copy the skill folder (.agents/skills/ai-evals in windmill-labs/windmill) into .claude/skills/ai-evals in your project. Claude Code loads it when a task matches its description.
Run `npx skills add windmill-labs/windmill --skill ai-evals -a codex`. Or copy the skill folder (.agents/skills/ai-evals in windmill-labs/windmill) into .agents/skills/ai-evals in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add windmill-labs/windmill --skill ai-evals -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ai-evals, .gemini/skills/ai-evals, .github/skills/ai-evals and .opencode/skills/ai-evals in your project.
Going by SKILL.md and its folder, Windmill AI Evals needs the command-line tools its instructions call (bun). Our summary lists: Bun and a Windmill checkout with the ai_evals directory; A reachable Windmill backend and an existing workspace; Model provider keys in ai_evals/.env.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.
Windmill AI Evals has a licence file (the repository's licence) that doesn't match a standard licence. Read it on GitHub before reusing the skill.
About 969 tokens (SKILL.md is roughly 3.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Windmill AI Evals: Agent Eval Engineering (langchain-ai/langchain-skills, 1.3k stars), Octocode Benchmark Runner (bgauryy/octocode, 949 stars), SWE Benchmark Task Adder (ory/lumen, 307 stars) and GAIA Agent Benchmarking (amd/gaia, 1.6k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
windmill-labs (a GitHub organization) maintains it in windmill-labs/windmill, which has 18,146 GitHub stars. The repository holds 42 skills in this directory. The repository was last updated on October 9, 2026.
Source: windmill-labs/windmill on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.