Agent Eval Engineering
langchain-ai/langchain-skills
Builds agent evaluations in stages: inspect the repository and traces, agree a Task Spec with you, then build, audit and run a Harbor task with an independent verifier.
Run skill evals via evals/runner.js — wrapper validates skill names, captures stdout/stderr, persists JSON results.
$ npx skills add ww-w-ai/bkit-claude-code --skill bkit-evals -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install ww-w-ai/bkit-claude-code bkit-evals --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/ww-w-ai/bkit-claude-code.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/bkit-evals .claude/skills/bkit-evals && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "bkit-evals" agent skill from https://github.com/ww-w-ai/bkit-claude-code/tree/main/skills/bkit-evals into .claude/skills/bkit-evals/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bkit-evals", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/ww-w-ai/bkit-claude-code/tree/main/skills/bkit-evalsType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add ww-w-ai/bkit-claude-code --skill bkit-evals -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install ww-w-ai/bkit-claude-code bkit-evals --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ww-w-ai/bkit-claude-code.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/bkit-evals .agents/skills/bkit-evals && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "bkit-evals" agent skill from https://github.com/ww-w-ai/bkit-claude-code/tree/main/skills/bkit-evals into .agents/skills/bkit-evals/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bkit-evals", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ww-w-ai/bkit-claude-code --skill bkit-evals -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install ww-w-ai/bkit-claude-code bkit-evals --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ww-w-ai/bkit-claude-code.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/bkit-evals .cursor/skills/bkit-evals && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "bkit-evals" agent skill from https://github.com/ww-w-ai/bkit-claude-code/tree/main/skills/bkit-evals into .cursor/skills/bkit-evals/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bkit-evals", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/ww-w-ai/bkit-claude-code.git --path skills/bkit-evals--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add ww-w-ai/bkit-claude-code --skill bkit-evals -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install ww-w-ai/bkit-claude-code bkit-evals --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ww-w-ai/bkit-claude-code.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/bkit-evals .gemini/skills/bkit-evals && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "bkit-evals" agent skill from https://github.com/ww-w-ai/bkit-claude-code/tree/main/skills/bkit-evals into .gemini/skills/bkit-evals/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bkit-evals", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install ww-w-ai/bkit-claude-code bkit-evalsInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add ww-w-ai/bkit-claude-code --skill bkit-evals -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/ww-w-ai/bkit-claude-code.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/bkit-evals .github/skills/bkit-evals && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "bkit-evals" agent skill from https://github.com/ww-w-ai/bkit-claude-code/tree/main/skills/bkit-evals into .github/skills/bkit-evals/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bkit-evals", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ww-w-ai/bkit-claude-code --skill bkit-evals -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install ww-w-ai/bkit-claude-code bkit-evals --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ww-w-ai/bkit-claude-code.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/bkit-evals .opencode/skills/bkit-evals && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "bkit-evals" agent skill from https://github.com/ww-w-ai/bkit-claude-code/tree/main/skills/bkit-evals into .opencode/skills/bkit-evals/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bkit-evals", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
bkit-evalsRun skill evals via evals/runner.js — wrapper validates skill names, captures stdout/stderr, persists JSON results.
Bkit Evals is an agent skill from ww-w-ai/bkit-claude-code. Run skill evals via evals/runner.js — wrapper validates skill names, captures stdout/stderr, persists JSON results. Triggers: bkit evals, evals run, skill quality, eval runner
Its SKILL.md is about 1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in AI & LLM Engineering, covering LLM evaluation and Agent evaluation and testing. The repository describes itself as: bkit Vibecoding Kit - PDCA methodology + Claude Code mastery for AI-native development. The licence is Apache-2.0.
6 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 85b4913. It shows what the files ask for, not the result of running them.
Pre-approves these tools, so the agent can use them without asking each time:
BashReadGlobGrepFrom allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
nodeFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Bkit Evals loads about 1k tokens when it runs. Until then it costs about 47 tokens; SKILL.md has 373 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
allowed-tools: Bash, Read, Glob, GrepAutomated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from ww-w-ai/bkit-claude-code at commit 85b4913, republished under its Apache-2.0 licence (© ww-w-ai). 373 words, ~1,015 tokens.
.claude/skills/bkit-evals/SKILL.md (or your agent's skills folder).v2.1.11 Sprint β FR-β2. Wraps
evals/runner.jswith input validation, result persistence, and structured reporting. Replaces the barenode evals/runner.js <skill>invocation that previously required users to remember argv structure and ignored timeout / sandbox concerns.
| Argument | Description | Example |
|---|---|---|
run <skill> | Execute the eval suite for one skill | /bkit-evals run gap-detector |
list | List all skills that have an eval.yaml definition | /bkit-evals list |
If no argument is provided, render the same output as list.
run <skill>skill against /^[a-z][a-z0-9-]{0,63}$/. Reject anything else
(no shell metacharacters, no slashes, no spaces) — see Security below.node evals/runner.js --skill <skill> via child_process.spawnSync
(argv form, no shell). Default timeout 30 s, max 120 s. The --skill flag
form is mandated by the runner CLI and locked by L3 contract test.parsed === null and stdout includes
Usage:, return reason: 'argv_format_mismatch'; if parsed === null
otherwise, return reason: 'parsed_null'. Exit code 0 alone NEVER
implies success — the parsed JSON must be present..bkit/runtime/evals-{skill}-{ISO timestamp}.json with stdout/stderr
tails (2000 chars each), parsed payload, and reason field.listevals/config.json to enumerate skill classifications.workflow, capability, hybrid),
list skills that have evals/{classification}/{skill}/eval.yaml.description field if present).[a-z][a-z0-9-]{0,63} is rejected with reason: invalid_skill_name.| Module | Function | Usage |
|---|---|---|
lib/evals/runner-wrapper.js | invokeEvals(skill, opts) | Validate + spawn + persist |
lib/evals/runner-wrapper.js | isValidSkillName(name) | Regex pre-check shared with list |
evals/runner.js | (subprocess) | Existing eval execution engine |
.bkit/runtime/evals-{skill}-{timestamp}.json:
{
"skill": "gap-detector",
"invokedAt": "<ISO 8601>",
"exitCode": 0,
"timedOut": false,
"stdoutTail": "...",
"stderrTail": "...",
"parsed": { /* whatever runner.js prints as JSON, or null */ }
}# Single eval
/bkit-evals run gap-detector
# Discovery
/bkit-evals list/control trust — eval results contribute to trust score/code-review — uses eval data when assessing skills/bkit explore (FR-β1) — explore evals as a categoryARGUMENTS:
© ww-w-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/bkit-evals of ww-w-ai/bkit-claude-code.
Open the folder on GitHubat commit 85b4913
Bkit Evals next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Bkit Evals this skillww-w-ai/bkit-claude-code | 601 | — | ~1k | Automated safety check: Notes | Apache-2.0 | |
| Agent Eval Engineeringlangchain-ai/langchain-skills | 1.3k | — | ~4k | Automated safety check: Pass | MIT | |
| GAIA Agent Benchmarkingamd/gaia | 1.6k | — | ~1.8k | Automated safety check: Pass | MIT | |
| Benchflowbenchflow-ai/benchflow | 356 | — | ~1.9k | Automated safety check: Notes | Apache-2.0 | |
| Email Evalstokencanopy/e2a | 193 | — | ~2.1k | Automated safety check: Pass | Apache-2.0 | |
| Windmill AI Evalswindmill-labs/windmill | 18k | — | ~969 | Automated safety check: Notes | Custom licence |
langchain-ai/langchain-skills
Builds agent evaluations in stages: inspect the repository and traces, agree a Task Spec with you, then build, audit and run a Harbor task with an independent verifier.
amd/gaia
Benchmarks AMD's GAIA agent against Claude Code and across models on quality, honesty, steps, tokens, time and real cost, using gaia eval tasks.
benchflow-ai/benchflow
Run agent benchmarks, create tasks, analyze results, and manage agents using BenchFlow.
tokencanopy/e2a
Author and safely run deterministic email-agent evaluation suites with dedicated e2a test agents.
windmill-labs/windmill
Writes and runs black-box benchmark cases for Windmill's flow, app, script, CLI and global AI generation modes, including before-and-after comparisons.
bgauryy/octocode
Runs blind pairwise comparisons of Octocode against a gh-based baseline over markdown research questions, scored by total characters through the model rather than self-report.
ww-w-ai/bkit-claude-code
View audit logs, decision traces, and session history for AI transparency.
ww-w-ai/bkit-claude-code
bkend.ai authentication — email/social login, JWT tokens, RBAC, session management.
ww-w-ai/bkit-claude-code
bkend.ai project tutorials (todo to SaaS) and common error troubleshooting.
ww-w-ai/bkit-claude-code
bkend.ai onboarding — MCP setup, resource hierarchy, tenant/user model, first project.
ww-w-ai/bkit-claude-code
bkend.ai file storage — upload (presigned URL), download (CDN), visibility levels, buckets.
ww-w-ai/bkit-claude-code
bkit plugin help - list available functions including /pdca (9-phase feature cycle), /sprint (8-phase feature container, v2.1.13), /control (Trust L0-L4 + SPRINTAUTORUNSCOPE), /bkit-explore, and 40+…
Categories
Run skill evals via evals/runner.js — wrapper validates skill names, captures stdout/stderr, persists JSON results. Bkit Evals is an agent skill from ww-w-ai/bkit-claude-code.js — wrapper validates skill names, captures stdout/stderr, persists JSON results.
Bkit Evals fits situations like: tasks that involve LLM evaluation; tasks that involve Agent evaluation and testing.
Run `npx skills add ww-w-ai/bkit-claude-code --skill bkit-evals -a claude-code`. Or copy the skill folder (skills/bkit-evals in ww-w-ai/bkit-claude-code) into .claude/skills/bkit-evals in your project. Claude Code loads it when a task matches its description.
Run `npx skills add ww-w-ai/bkit-claude-code --skill bkit-evals -a codex`. Or copy the skill folder (skills/bkit-evals in ww-w-ai/bkit-claude-code) into .agents/skills/bkit-evals in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ww-w-ai/bkit-claude-code --skill bkit-evals -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bkit-evals, .gemini/skills/bkit-evals, .github/skills/bkit-evals and .opencode/skills/bkit-evals in your project.
Going by SKILL.md and its folder, Bkit Evals needs the command-line tools its instructions call (node). Its frontmatter pre-approves these tools: Bash, Read, Glob, Grep.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.
Bkit Evals is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1k tokens (SKILL.md is roughly 4.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Bkit Evals: Agent Eval Engineering (langchain-ai/langchain-skills, 1.3k stars), GAIA Agent Benchmarking (amd/gaia, 1.6k stars), Benchflow (benchflow-ai/benchflow, 356 stars) and Email Evals (tokencanopy/e2a, 193 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
ww-w-ai (a GitHub organization) maintains it in ww-w-ai/bkit-claude-code, which has 601 GitHub stars. The repository holds 44 skills in this directory. The repository was last updated on September 27, 2026.
Source: ww-w-ai/bkit-claude-code on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.