Agent skill

Bkit Evals

by ww-w-ai in ww-w-ai/bkit-claude-code

Run skill evals via evals/runner.js — wrapper validates skill names, captures stdout/stderr, persists JSON results.

Apache-2.0Auto-check: notesAI & LLM Engineering

Install Bkit Evals

skills CLI
$ npx skills add ww-w-ai/bkit-claude-code --skill bkit-evals -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ww-w-ai/bkit-claude-code bkit-evals --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ww-w-ai/bkit-claude-code.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/bkit-evals .claude/skills/bkit-evals && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bkit-evals
GitHub stars
601
Token cost
~1k tokens
SKILL.md length
373 words
Files
1
Skills in repo
44
Repo updated
First seen
Licence
Apache-2.0

At a glance

Run skill evals via evals/runner.js — wrapper validates skill names, captures stdout/stderr, persists JSON results.

  • Works in 6 steps: Validate skill against /^a-z{0,63}$/.… → Spawn node evals/runner.js --skill via… → Capture stdout / stderr. Parse the… → …
  • Tasks that involve LLM evaluation
  • SKILL.md covers Arguments, Behavior, Security and Module Dependencies, plus 3 more sections
  • Calls node

What it does

Bkit Evals is an agent skill from ww-w-ai/bkit-claude-code. Run skill evals via evals/runner.js — wrapper validates skill names, captures stdout/stderr, persists JSON results. Triggers: bkit evals, evals run, skill quality, eval runner

Its SKILL.md is about 1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering LLM evaluation and Agent evaluation and testing. The repository describes itself as: bkit Vibecoding Kit - PDCA methodology + Claude Code mastery for AI-native development. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve LLM evaluation
  • Tasks that involve Agent evaluation and testing

Example prompts

  • “/bkit-evals”

Requirements

  • Pre-approved tools (allowed-tools): Bash, Read, Glob, Grep

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Validate skill against /^a-z{0,63}$/. Reject anything else
  2. Spawn node evals/runner.js --skill via child_process.spawnSync
  3. Capture stdout / stderr. Parse the trailing JSON block via
  4. Apply fail-closed defense: if parsed === null and stdout includes
  5. Persist the structured result to
  6. Render a one-line summary in the chat

What it can do on your machine

Read from SKILL.md and the folder at commit 85b4913. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Bash
    • Read
    • Glob
    • Grep

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • node

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bkit Evals loads about 1k tokens when it runs. Until then it costs about 47 tokens; SKILL.md has 373 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~47
When it runs · the whole SKILL.md, loaded when a task matches
~1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Bash, Read, Glob, Grep

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from ww-w-ai/bkit-claude-code at commit 85b4913, republished under its Apache-2.0 licence (© ww-w-ai). 373 words, ~1,015 tokens.

Download SKILL.mdSave it as .claude/skills/bkit-evals/SKILL.md (or your agent's skills folder).
name
bkit-evals
description
Run skill evals via evals/runner.js — wrapper validates skill names, captures stdout/stderr, persists JSON results. Triggers: bkit evals, evals run, skill quality, eval runner
allowed-tools
Bash, Read, Glob, Grep
classification
capability
classification-reason
Eval runner is a development-time quality tool, not a workflow phase
deprecation-risk
none
effort
low
argument-hint
run <skill> | list
user-invocable
true
task-template
[Evals] {action}

bkit Evals — Skill Quality Evaluation Runner

v2.1.11 Sprint β FR-β2. Wraps evals/runner.js with input validation, result persistence, and structured reporting. Replaces the bare node evals/runner.js <skill> invocation that previously required users to remember argv structure and ignored timeout / sandbox concerns.

Arguments

ArgumentDescriptionExample
run <skill>Execute the eval suite for one skill/bkit-evals run gap-detector
listList all skills that have an eval.yaml definition/bkit-evals list

If no argument is provided, render the same output as list.

Behavior

run <skill>
  1. Validate skill against /^[a-z][a-z0-9-]{0,63}$/. Reject anything else (no shell metacharacters, no slashes, no spaces) — see Security below.
  2. Spawn node evals/runner.js --skill <skill> via child_process.spawnSync (argv form, no shell). Default timeout 30 s, max 120 s. The --skill flag form is mandated by the runner CLI and locked by L3 contract test.
  3. Capture stdout / stderr. Parse the trailing JSON block via balanced-brace fallback (string-aware).
  4. Apply fail-closed defense: if parsed === null and stdout includes Usage:, return reason: 'argv_format_mismatch'; if parsed === null otherwise, return reason: 'parsed_null'. Exit code 0 alone NEVER implies success — the parsed JSON must be present.
  5. Persist the structured result to .bkit/runtime/evals-{skill}-{ISO timestamp}.json with stdout/stderr tails (2000 chars each), parsed payload, and reason field.
  6. Render a one-line summary in the chat:
    • exit code
    • parsed pass/fail counts (if available)
    • path of the persisted result file
Show full SKILL.md (152 more words)Show less
list
  1. Read evals/config.json to enumerate skill classifications.
  2. For each classification (workflow, capability, hybrid), list skills that have evals/{classification}/{skill}/eval.yaml.
  3. Render a category-grouped table with skill name + a one-line note from the eval YAML (description field if present).

Security

  • Skill name regex prevents argument injection. Anything outside [a-z][a-z0-9-]{0,63} is rejected with reason: invalid_skill_name.
  • argv-array spawn (no shell). No template-string concatenation into command lines.
  • Result file path is composed from a hardcoded base + sanitized skill name + timestamp; no traversal possible.
  • Subprocess timeout enforced (default 30 s, hard cap 120 s) so a buggy eval cannot block the session indefinitely.

Module Dependencies

ModuleFunctionUsage
lib/evals/runner-wrapper.jsinvokeEvals(skill, opts)Validate + spawn + persist
lib/evals/runner-wrapper.jsisValidSkillName(name)Regex pre-check shared with list
evals/runner.js(subprocess)Existing eval execution engine

Result Schema

.bkit/runtime/evals-{skill}-{timestamp}.json:

json
{
  "skill": "gap-detector",
  "invokedAt": "<ISO 8601>",
  "exitCode": 0,
  "timedOut": false,
  "stdoutTail": "...",
  "stderrTail": "...",
  "parsed": { /* whatever runner.js prints as JSON, or null */ }
}

Examples

bash
# Single eval
/bkit-evals run gap-detector

# Discovery
/bkit-evals list
  • /control trust — eval results contribute to trust score
  • /code-review — uses eval data when assessing skills
  • /bkit explore (FR-β1) — explore evals as a category

ARGUMENTS:

© ww-w-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/bkit-evals of ww-w-ai/bkit-claude-code.

Open the folder on GitHubat commit 85b4913

Compare with similar skills

Bkit Evals next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bkit Evals compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bkit Evals this skillww-w-ai/bkit-claude-code601—~1kAutomated safety check: NotesApache-2.0
Agent Eval Engineeringlangchain-ai/langchain-skills1.3k—~4kAutomated safety check: PassMIT
GAIA Agent Benchmarkingamd/gaia1.6k—~1.8kAutomated safety check: PassMIT
Benchflowbenchflow-ai/benchflow356—~1.9kAutomated safety check: NotesApache-2.0
Email Evalstokencanopy/e2a193—~2.1kAutomated safety check: PassApache-2.0
Windmill AI Evalswindmill-labs/windmill18k—~969Automated safety check: NotesCustom licence

Similar skills

  • Agent Eval Engineering

    langchain-ai/langchain-skills

    Official

    Builds agent evaluations in stages: inspect the repository and traces, agree a Task Spec with you, then build, audit and run a Harbor task with an independent verifier.

    1.3k GitHub stars~4k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Benchmarks AMD's GAIA agent against Claude Code and across models on quality, honesty, steps, tokens, time and real cost, using gaia eval tasks.

    1.6k GitHub stars~1.8k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Benchflow

    benchflow-ai/benchflow

    Run agent benchmarks, create tasks, analyze results, and manage agents using BenchFlow.

    356 GitHub stars~1.9k tokensUpdated 5 days ago
    AI & LLM EngineeringAuto-check: notes
  • Email Evals

    tokencanopy/e2a

    Author and safely run deterministic email-agent evaluation suites with dedicated e2a test agents.

    193 GitHub stars~2.1k tokensUpdated 4 days ago
    AI & LLM EngineeringAuto-check passed
  • Windmill AI Evals

    windmill-labs/windmill

    Writes and runs black-box benchmark cases for Windmill's flow, app, script, CLI and global AI generation modes, including before-and-after comparisons.

    18k GitHub stars~969 tokensUpdated yesterday
    AI & LLM EngineeringAuto-check: notes
  • Runs blind pairwise comparisons of Octocode against a gh-based baseline over markdown research questions, scored by total characters through the model rather than self-report.

    949 GitHub stars~2.1k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed

More from ww-w-ai/bkit-claude-code

All 44 skills in this repo
  • Audit

    ww-w-ai/bkit-claude-code

    View audit logs, decision traces, and session history for AI transparency.

    601 GitHub stars~1.6k tokensUpdated 13 days ago
    Auto-check: notes
  • Bkend Auth

    ww-w-ai/bkit-claude-code

    bkend.ai authentication — email/social login, JWT tokens, RBAC, session management.

    601 GitHub stars~937 tokensUpdated 13 days ago
    Auto-check: notes
  • Bkend Cookbook

    ww-w-ai/bkit-claude-code

    bkend.ai project tutorials (todo to SaaS) and common error troubleshooting.

    601 GitHub stars~891 tokensUpdated 13 days ago
    Auto-check: notes
  • Bkend Quickstart

    ww-w-ai/bkit-claude-code

    bkend.ai onboarding — MCP setup, resource hierarchy, tenant/user model, first project.

    601 GitHub stars~1.2k tokensUpdated 13 days ago
    Auto-check passed
  • Bkend Storage

    ww-w-ai/bkit-claude-code

    bkend.ai file storage — upload (presigned URL), download (CDN), visibility levels, buckets.

    601 GitHub stars~901 tokensUpdated 13 days ago
    Auto-check: notes
  • Bkit

    ww-w-ai/bkit-claude-code

    bkit plugin help - list available functions including /pdca (9-phase feature cycle), /sprint (8-phase feature container, v2.1.13), /control (Trust L0-L4 + SPRINTAUTORUNSCOPE), /bkit-explore, and 40+…

    601 GitHub stars~1.4k tokensUpdated 13 days ago
    Auto-check passed

Questions about Bkit Evals

What does Bkit Evals do?

Run skill evals via evals/runner.js — wrapper validates skill names, captures stdout/stderr, persists JSON results. Bkit Evals is an agent skill from ww-w-ai/bkit-claude-code.js — wrapper validates skill names, captures stdout/stderr, persists JSON results.

When should I use Bkit Evals?

Bkit Evals fits situations like: tasks that involve LLM evaluation; tasks that involve Agent evaluation and testing.

How do I install Bkit Evals in Claude Code?

Run `npx skills add ww-w-ai/bkit-claude-code --skill bkit-evals -a claude-code`. Or copy the skill folder (skills/bkit-evals in ww-w-ai/bkit-claude-code) into .claude/skills/bkit-evals in your project. Claude Code loads it when a task matches its description.

How do I install Bkit Evals in Codex?

Run `npx skills add ww-w-ai/bkit-claude-code --skill bkit-evals -a codex`. Or copy the skill folder (skills/bkit-evals in ww-w-ai/bkit-claude-code) into .agents/skills/bkit-evals in your project. Codex loads it when a task matches its description.

Can I use Bkit Evals in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ww-w-ai/bkit-claude-code --skill bkit-evals -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bkit-evals, .gemini/skills/bkit-evals, .github/skills/bkit-evals and .opencode/skills/bkit-evals in your project.

What does Bkit Evals need to run?

Going by SKILL.md and its folder, Bkit Evals needs the command-line tools its instructions call (node). Its frontmatter pre-approves these tools: Bash, Read, Glob, Grep.

Does Bkit Evals access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Bkit Evals safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Bkit Evals use?

Bkit Evals is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bkit Evals use?

About 1k tokens (SKILL.md is roughly 4.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bkit Evals?

Skills that share tags, products or a category with Bkit Evals: Agent Eval Engineering (langchain-ai/langchain-skills, 1.3k stars), GAIA Agent Benchmarking (amd/gaia, 1.6k stars), Benchflow (benchflow-ai/benchflow, 356 stars) and Email Evals (tokencanopy/e2a, 193 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bkit Evals?

ww-w-ai (a GitHub organization) maintains it in ww-w-ai/bkit-claude-code, which has 601 GitHub stars. The repository holds 44 skills in this directory. The repository was last updated on September 27, 2026.

Source: ww-w-ai/bkit-claude-code on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.