Agent skill

Bat Story Eval

by homeassistant-ai in homeassistant-ai/ha-mcp

Compare MCP tool behavior between target and baseline versions using pre-built and custom stories with diff-based triage.

MITAuto-check: notesAgent Workflows

Install Bat Story Eval

skills CLI
$ npx skills add homeassistant-ai/ha-mcp --skill bat-story-eval -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install homeassistant-ai/ha-mcp bat-story-eval --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/homeassistant-ai/ha-mcp.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/bat-story-eval .claude/skills/bat-story-eval && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bat-story-eval
GitHub stars
5k
Token cost
~3.4k tokens
SKILL.md length
1,079 words
Files
3 (incl. references)
Skills in repo
8
Repo updated
First seen
Licence
MIT

At a glance

Compare MCP tool behavior between target and baseline versions using pre-built and custom stories with diff-based triage.

  • Works in 9 steps: Triage (Diff Analysis + Custom Story… → Run Baseline Version → Run Target Version → …
  • Tasks that involve MCP servers
  • SKILL.md covers Parse Arguments, Step 0: Triage (Diff Analysis…, Step 1: Run Baseline Version and Step 2: Run Target Version, plus 8 more sections
  • Calls uv, git and docker

What it does

Bat Story Eval is an agent skill from homeassistant-ai/ha-mcp. Compare MCP tool behavior between target and baseline versions using pre-built and custom stories with diff-based triage.

Its SKILL.md is about 3.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `references/evaluation-protocol.md` and `references/regression-protocol.md`).

It sits in Agent Workflows, covering MCP servers. The repository describes itself as: The Unofficial and Awesome Home Assistant MCP Server. The licence is MIT.

When your agent uses it

  • Tasks that involve MCP servers

Example prompts

  • “/bat-story-eval”

Requirements

  • Python 3
  • Docker
  • Pre-approved tools (allowed-tools): Bash, Read, Write, Glob, Grep, Task

Workflow steps

9 steps, taken from the step headings in SKILL.md.

  1. Triage (Diff Analysis + Custom Story Design)
  2. Run Baseline Version
  3. Run Target Version
  4. White-Box Analysis
  5. Score & Compare
  6. Update JSONL
  7. Report
  8. Investigate Outliers
  9. Tool Description Size

What it can do on your machine

Read from SKILL.md and the folder at commit b274a93. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Bash
    • Read
    • Write
    • Glob
    • Grep
    • Task

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • uv
    • git
    • docker
    • python3
    • claude

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uv, git and docker, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bat Story Eval loads about 3.4k tokens when it runs, and up to ~6.3k if it reads all its reference files. Until then it costs about 34 tokens; SKILL.md has 1,079 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~34
When it runs · the whole SKILL.md, loaded when a task matches
~3.4k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~6.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Bash, Read, Write, Glob, Grep, Task

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from homeassistant-ai/ha-mcp at commit b274a93, republished under its MIT licence (© homeassistant-ai). 1,079 words, ~3,398 tokens.

Download SKILL.mdSave it as .claude/skills/bat-story-eval/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
bat-story-eval
description
Compare MCP tool behavior between target and baseline versions using pre-built and custom stories with diff-based triage.
allowed-tools
Bash, Read, Write, Glob, Grep, Task
disable-model-invocation
true
argument-hint
--baseline v6.6.1 [--agents gemini] [--stories s01,s02]

BAT Story Evaluation

You are the evaluator. Run the steps in order; each one uses the output of the step before it.

Parse Arguments

From $ARGUMENTS, extract:

  • --baseline: REQUIRED. Git tag/branch of the released version (e.g., v6.6.1).
  • --agents: Agent list (default: gemini). Comma-separated.
  • --stories: Force specific pre-built story IDs (e.g., s01,s02). Overrides triage selection.
  • --all-stories: Skip triage, run ALL pre-built stories.
  • --keep-container: Keep HA containers alive after run for manual inspection.
  • --model: Model for Claude agent (e.g., haiku, sonnet).

If $ARGUMENTS is --help or missing --baseline, show usage and stop:

/bat-story-eval --baseline v6.6.1
/bat-story-eval --baseline v6.6.1 --agents gemini,claude
/bat-story-eval --baseline v6.6.1 --stories s01,s02
/bat-story-eval --baseline v6.6.1 --all-stories --agents claude --model haiku

Step 0: Triage (Diff Analysis + Custom Story Design)

0a. Compute Diff
bash
cd "$(dirname "$(git rev-parse --path-format=absolute --git-common-dir)")/worktree/uat-stories"
git diff <baseline>..HEAD -- src/ha_mcp/ --stat
git diff <baseline>..HEAD -- src/ha_mcp/ --name-only

Classify changed files:

  • Tool modules (tools/tools_*.py): specific tool implementations changed
  • Core code (client/, server.py, errors.py, tools/util_helpers.py): affects all tools
  • Utilities (utils/, resources/): may affect all tools
  • No src/ changes: only tests/docs/config — select 2 smoke-test stories
0b. Select Pre-built Stories

Skip if --stories or --all-stories was passed.

  1. Read the diff from 0a
  2. Read all story YAMLs in tests/uat/stories/catalog/s*.yaml (title, description, prompt, setup)
  3. For each story, reason about whether the diff could affect its outcome:
    • What tools/code paths would this story exercise?
    • Do any of those overlap with what changed?
  4. Rules:
    • Story likely exercises changed code -> selected
    • Core code changed (client/, server.py, errors.py) -> all stories selected
    • No src/ changes -> 2 representative stories as smoke test
  5. Report which stories were selected and why (one sentence per story)
0c. Design Custom Stories (at least 1)

Read the diff carefully. Your job is to catch regressions. For each changed code path NOT covered by selected pre-built stories, ask: "could this break something a user would notice?" If yes, design a custom story to test that hypothesis.

Guidelines: Always create at least 1 custom story. Each must test a distinct regression hypothesis — don't create stories that overlap. Stop when you've covered the risky gaps.

Write each as /tmp/custom_c<NN>.yaml using the standard story format:

yaml
id: c01
title: "Short description of what is being tested"
category: custom
weight: 5
description: >
  Rationale: [what changed in the diff and why this scenario tests it]

setup:
  - tool: ha_config_set_helper
    args:
      helper_type: "input_boolean"
      name: "Test Entity Name"

prompt: >
  [Natural language request a real user would make that exercises the changed code]

teardown: []

verify:
  questions:
    - "Did the agent achieve the expected outcome?"
    - "Did it use the expected tools?"

expected:
  tools_should_use:
    - ha_search
  description: >
    [What a correct agent should do]

Design principles:

  • Focus on code paths that changed in the diff
  • Plausible user scenarios, not synthetic edge cases
  • Setup creates realistic HA state via FastMCP in-memory steps
  • Prompts are what a real user would type
  • At least 1. Each tests a distinct regression hypothesis. Stop when gaps are covered.

Step 1: Run Baseline Version

For EACH agent, run all stories against the baseline version. One container per agent, reused across all stories.

1a. Start container with first story
bash
cd "$(dirname "$(git rev-parse --path-format=absolute --git-common-dir)")/worktree/uat-stories"
uv run python tests/uat/stories/run_story.py \
  tests/uat/stories/catalog/<first_story>.yaml \
  --agents <agent> --keep-container \
  --branch <baseline> \
  --results-file local/uat-results.jsonl

CAPTURE from stderr: HA URL (e.g., http://localhost:32771), token, session file path.

1b. Verify, then run remaining pre-built stories

After each story, verify via ha_query.py using the story's verify.questions:

bash
uv run python tests/uat/stories/scripts/ha_query.py \
  --ha-url http://localhost:PORT --ha-token TOKEN \
  --agent <agent> \
  "Does an automation with alias 'Sunset Porch Light' exist?"

Record each answer as confirmed / denied / unclear. A non-zero exit from ha_query.py means the query itself failed (the output carries an [exit N] marker; [exit 124] is a timeout) — that is not one of the three outcomes; re-run it, and if it keeps failing record the story as unverified (Step 5) rather than scoring it. See references/evaluation-protocol.md.

Run remaining pre-built stories on the same container:

bash
uv run python tests/uat/stories/run_story.py \
  tests/uat/stories/catalog/<next_story>.yaml \
  --agents <agent> --ha-url http://localhost:PORT --ha-token TOKEN \
  --branch <baseline> \
  --results-file local/uat-results.jsonl

Verify each immediately after running.

1c. Run custom stories on same container
bash
uv run python tests/uat/stories/run_story.py \
  /tmp/custom_c01.yaml \
  --agents <agent> --ha-url http://localhost:PORT --ha-token TOKEN \
  --branch <baseline> \
  --results-file local/uat-results.jsonl

Verify each via ha_query.py using the custom story's verify.questions.

1d. Stop container

Stop only the container kept in step 1a. PORT is the host port in its Container kept alive: http://localhost:PORT line:

bash
docker stop $(docker ps -q --filter "publish=PORT")

Step 2: Run Target Version

Repeat Step 1 for the target (local code). Same stories, same order, fresh container.

The only difference: omit --branch so run_story.py uses local code.

bash
uv run python tests/uat/stories/run_story.py \
  tests/uat/stories/catalog/<first_story>.yaml \
  --agents <agent> --keep-container \
  --results-file local/uat-results.jsonl

Same container reuse for remaining stories (--ha-url). Same verification after each.

Step 3: White-Box Analysis

For each story on each version, read the session file captured during the run.

Gemini sessions (JSON):

bash
python3 -c "
import json, sys
data = json.load(open(sys.argv[1]))
for msg in data.get('messages', []):
    for tc in msg.get('toolCalls', []):
        print(f\"  {tc['name']} ({tc.get('status', '?')})\")
" /path/to/session.json

Claude sessions (JSONL):

bash
python3 -c "
import json, sys
for line in open(sys.argv[1]):
    entry = json.loads(line)
    if entry.get('type') == 'assistant':
        for b in entry.get('message', {}).get('content', []):
            if b.get('type') == 'tool_use':
                print(f\"  {b['name']}\")
" /path/to/session.jsonl

Compare against expected.tools_should_use:

  • All expected tools used? (High weight)
  • Tool failures with recovery? (Medium weight)
  • Total tool call count (Low weight, note it)

Step 4: Score & Compare

Scoring Matrix
Black-BoxWhite-BoxScore
Entity correct + right structureRight toolspass
Entity correct + right structureWrong tools or recovered errorspass (with notes)
Entity correct + wrong structureAnypartial
Entity not createdAnyfail
Show full SKILL.md (425 more words)Show less
Metrics

Primary metrics (decide pass/fail on these):

  • Black-box score (entity correct, structure correct)
  • White-box tool selection (expected tools used)
  • Error recovery (failures handled gracefully)

Secondary metrics (report but don't decide on these alone):

  • Billable tokens — directional cost signal, flag >30% increase for investigation but don't auto-fail
  • Cached tokens / cache hit ratio — useful context for cost analysis, but varies based on provider-side KV-cache behavior
  • Tool call count / turns — varies between runs due to agent exploration
  • Duration — noisy (network, KV-cache misses, server load), only flag large (>2x) outliers
  • Tool description size delta (Step 8)
Extracting Billable Tokens
python
# Gemini: input includes cached, so subtract
billable = (input - cached) + output + thoughts

# Claude: input_tokens is already non-cached
billable = input + output
Trend (target vs baseline)

For each story+agent:

  • Both pass -> stable
  • Target pass, baseline fail -> improved
  • Target fail, baseline pass -> decreased (REGRESSION)
  • Custom story, first run -> new
  • Billable tokens >30% higher -> cost investigation (even if pass — check Step 7 for KV-cache misses before concluding regression)

Step 5: Update JSONL

Append eval results as NEW lines (never modify existing):

python
record["eval_score"] = "pass"  # or "partial", "fail", or "unverified"
record["eval_notes"] = "Entity created, triggers verified"
record["eval_trend"] = "stable"  # or "new", "improved", "decreased"
# "unverified" is for a story whose verification query itself failed — it is
# not a result, so it carries no trend and is not compared to the baseline.

Step 6: Report

Triage Summary
Diff: <baseline>..HEAD — N files changed in src/ha_mcp/
Selected pre-built: s01, s03, s05 (3 stories — tools_automation.py, tools_search.py changed)
Custom stories: c01, c02 (2 stories — covering error handling, fuzzy search threshold)
Skipped: s02, s04, s06-s12 (tools unchanged)
Pre-built Story Results

Also read the model and quantization fields from each JSONL record (written by run_story.py) and show them as columns. Results vary by model, and the same base model at different quants behaves very differently, so a report naming only the agent is ambiguous after the fact. (Quant is - for cloud backends that don't expose it.)

| Story | Agent  | Model            | Quant | Baseline | Target | Trend  | Baseline Tokens | Target Tokens | Delta |
|-------|--------|------------------|-------|----------|--------|--------|-----------------|---------------|-------|
| s01   | claude | claude-sonnet-4-6| -     | pass     | pass   | stable | 36,262          | 34,100        | -6%   |
| s03   | claude | claude-sonnet-4-6| -     | pass     | pass   | stable | 42,000          | 41,500        | -1%   |
Custom Story Details

For EACH custom story, output a full section:

#### c01: [Title]

**Rationale**: [What changed in the diff and why this tests it]

**Setup**:
- Created input_boolean "Sophisticated Kitchen Sensor" via FastMCP

**Test prompt**: "[The exact prompt sent to the agent]"

**Verification**:
| Question | Baseline | Target |
|----------|----------|--------|
| Found the entity? | confirmed | confirmed |
| Used ha_search? | confirmed | confirmed |

**Score**: baseline=pass, target=pass, trend=stable
**Tokens**: baseline=28,500, target=27,200 (-5%)
Regressions

If any trend = decreased:

  1. Flag prominently
  2. Suggest re-run to check flakiness
  3. Show relevant section of git diff <baseline>..HEAD

Step 7: Investigate Outliers

When a story has >30% more billable tokens vs baseline, check for KV-cache misses:

python
for i, msg in enumerate(data["messages"]):
    tok = msg.get("tokens", {})
    cached = tok.get("cached", 0)
    total = tok.get("input", 0)
    print(f"Turn {i+1}: input={total:,} cached={cached:,} non-cached={total-cached:,}")

A turn with cached=0 after a non-cold-start turn = KV-cache miss (provider-side, not a code regression).

Step 8: Tool Description Size

Compare tool description sizes between versions:

bash
uv run python tests/uat/stories/scripts/measure_tools.py \
  --output local/tool-sizes-target.json
uv run python tests/uat/stories/scripts/measure_tools.py \
  --output local/tool-sizes-baseline.json --branch <baseline>

Flag >5% total size increase (directly impacts token cost per turn).

Key Files

FilePurpose
tests/uat/stories/run_story.pyStory runner (container, setup, agent CLI)
tests/uat/stories/scripts/ha_query.pyQuery live HA via agent+MCP for verification
tests/uat/stories/catalog/s*.yamlPre-built story definitions
local/uat-results.jsonlHistorical results (gitignored)
references/evaluation-protocol.mdScoring rules, verification questions, cross-agent checks
references/regression-protocol.mdRegression classification and flaky handling

Important Notes

  • --baseline is required: it's both the diff source and the control group
  • Run pre-built stories BEFORE custom stories (cleanest state)
  • ALWAYS verify each story via ha_query.py before running the next
  • Reuse containers: first story starts it (--keep-container), rest use --ha-url
  • Custom story YAMLs go to /tmp/ (ephemeral); full details reported in Step 6
  • See "Metrics" section in Step 4 for primary vs secondary metric classification
  • The working directory MUST be the worktree/uat-stories worktree root (where pyproject.toml lives) for uv run

© homeassistant-ai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references) in .claude/skills/bat-story-eval of homeassistant-ai/ha-mcp.

  • SKILL.md
  • references/evaluation-protocol.md
  • references/regression-protocol.md

Open the folder on GitHubat commit b274a93

Compare with similar skills

Bat Story Eval next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bat Story Eval compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bat Story Eval this skillhomeassistant-ai/ha-mcp5k—~3.4kAutomated safety check: NotesMIT
MCP Server Builderanthropics/skills180k63 repos~2.3kAutomated safety check: PassApache-2.0
MCP Server BuildershareAI-lab/learn-claude-code78k4 repos~1.2kAutomated safety check: PassMIT
MCP Integration for Pluginsanthropics/claude-plugins-official38k11 repos~3.1kAutomated safety check: PassApache-2.0
Microsoft Skill CreatorMicrosoftDocs/mcp1.9k3 repos~2.1kAutomated safety check: PassCC-BY-4.0
Crush Configurationcharmbracelet/crush29k—~3.7kAutomated safety check: PassCustom licence

Similar skills

  • MCP Server Builder

    anthropics/skills

    Official

    Guides the design and implementation of Model Context Protocol servers in TypeScript or Python, from tool naming and error messages to evaluation.

    180k GitHub starsUsed in 63 repos~2.3k tokens
    Agent WorkflowsAuto-check passed
  • MCP Server Builder

    shareAI-lab/learn-claude-code

    Walks through building MCP servers in Python or TypeScript that expose tools, resources and prompts to Claude, with templates, registration and testing.

    78k GitHub starsUsed in 4 repos~1.2k tokens
    Agent WorkflowsAuto-check passed
  • MCP Integration for Plugins

    anthropics/claude-plugins-official

    Official

    Explains how to bundle Model Context Protocol servers in a Claude Code plugin, covering config files, stdio, SSE, HTTP and WebSocket server types, and authentication.

    38k GitHub starsUsed in 11 repos~3.1k tokens
    Agent WorkflowsAuto-check passed
  • Microsoft Skill Creator

    MicrosoftDocs/mcp

    Official

    Create agent skills for Microsoft technologies using official documentation.

    1.9k GitHub starsUsed in 3 repos~2.1k tokens
    Agent WorkflowsAuto-check passed
  • Crush Configuration

    charmbracelet/crush

    Explains how to configure the Crush coding agent with crushrc or crush.json, covering providers, models, LSPs, MCP servers, hooks, permissions and config precedence.

    29k GitHub stars~3.7k tokensUpdated today
    Agent WorkflowsAuto-check passed
  • Context Mode Output Sandbox

    mksglu/context-mode

    Routes large command, file, API and browser output through context-mode tools so only the needed result enters the agent's context, instead of dumping it via Bash.

    26k GitHub stars~4.1k tokensUpdated today
    Agent WorkflowsAuto-check passed

More from homeassistant-ai/ha-mcp

All 8 skills in this repo
  • Bat Adhoc

    homeassistant-ai/ha-mcp

    Run bot acceptance tests to validate MCP tools work correctly from a real AI agent's perspective.

    5k GitHub stars~1.4k tokensUpdated today
    Auto-check: notes
  • Contrib PR Review

    homeassistant-ai/ha-mcp

    Review a contribution PR for safety, quality, and readiness.

    5k GitHub stars~3.1k tokensUpdated today
    Auto-check: notes
  • Issue Analysis

    homeassistant-ai/ha-mcp

    Deep analysis of a single GitHub issue with codebase exploration, implementation planning, and architectural assessment.

    5k GitHub stars~753 tokensUpdated today
    Auto-check: notes
  • Issue To PR Resolver

    homeassistant-ai/ha-mcp

    Implement a GitHub issue end-to-end — create a worktree branch, implement the feature with tests, create a draft PR, then iteratively resolve all CI failures and review comments until the PR is clean.

    5k GitHub stars~1.2k tokensUpdated today
    Auto-check: notes
  • My PR Checker

    homeassistant-ai/ha-mcp

    Manage your own GitHub pull requests — check CI status, inline review comments, PR-level comments, resolve review threads, fix issues, and iterate until all checks pass and threads are resolved.

    5k GitHub stars~1.2k tokensUpdated today
    Auto-check: notes
  • Contributors Update

    homeassistant-ai/ha-mcp

    Find merged PR authors missing from README and update the contributors list after approval

    5k GitHub stars~983 tokensUpdated today
    Auto-check passed

Categories

Questions about Bat Story Eval

What does Bat Story Eval do?

Compare MCP tool behavior between target and baseline versions using pre-built and custom stories with diff-based triage. Bat Story Eval is an agent skill from homeassistant-ai/ha-mcp. Compare MCP tool behavior between target and baseline versions using pre-built and custom stories with diff-based triage.

When should I use Bat Story Eval?

Bat Story Eval fits situations like: tasks that involve MCP servers.

How do I install Bat Story Eval in Claude Code?

Run `npx skills add homeassistant-ai/ha-mcp --skill bat-story-eval -a claude-code`. Or copy the skill folder (.claude/skills/bat-story-eval in homeassistant-ai/ha-mcp) into .claude/skills/bat-story-eval in your project. Claude Code loads it when a task matches its description.

How do I install Bat Story Eval in Codex?

Run `npx skills add homeassistant-ai/ha-mcp --skill bat-story-eval -a codex`. Or copy the skill folder (.claude/skills/bat-story-eval in homeassistant-ai/ha-mcp) into .agents/skills/bat-story-eval in your project. Codex loads it when a task matches its description.

Can I use Bat Story Eval in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add homeassistant-ai/ha-mcp --skill bat-story-eval -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bat-story-eval, .gemini/skills/bat-story-eval, .github/skills/bat-story-eval and .opencode/skills/bat-story-eval in your project.

What does Bat Story Eval need to run?

Going by SKILL.md and its folder, Bat Story Eval needs the command-line tools its instructions call (uv, git, docker, python3 and claude). Our summary lists: Python 3; Docker. Its frontmatter pre-approves these tools: Bash, Read, Write, Glob, Grep, Task.

Does Bat Story Eval access the network?

SKILL.md contains no URLs. Its commands use uv, git and docker, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Bat Story Eval safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Bat Story Eval use?

Bat Story Eval is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bat Story Eval use?

About 3.4k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.9k tokens, read only when the agent opens those files.

What are the alternatives to Bat Story Eval?

Skills that share tags, products or a category with Bat Story Eval: MCP Server Builder (anthropics/skills, 180k stars), MCP Server Builder (shareAI-lab/learn-claude-code, 78k stars), MCP Integration for Plugins (anthropics/claude-plugins-official, 38k stars) and Microsoft Skill Creator (MicrosoftDocs/mcp, 1.9k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bat Story Eval?

homeassistant-ai (a GitHub organization) maintains it in homeassistant-ai/ha-mcp, which has 5,015 GitHub stars. The repository holds 8 skills in this directory. The repository was last updated on October 11, 2026.

Source: homeassistant-ai/ha-mcp on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.