Agent skill

Bat Adhoc

by homeassistant-ai in homeassistant-ai/ha-mcp

Run bot acceptance tests to validate MCP tools work correctly from a real AI agent's perspective.

MITAuto-check: notesTesting & QA

Install Bat Adhoc

skills CLI
$ npx skills add homeassistant-ai/ha-mcp --skill bat-adhoc -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install homeassistant-ai/ha-mcp bat-adhoc --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/homeassistant-ai/ha-mcp.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/bat-adhoc .claude/skills/bat-adhoc && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bat-adhoc
GitHub stars
5k
Token cost
~1.4k tokens
SKILL.md length
506 words
Files
1
Skills in repo
8
Repo updated
First seen
Licence
MIT

At a glance

Run bot acceptance tests to validate MCP tools work correctly from a real AI agent's perspective.

  • Works in 6 steps: Analyze the change: Read the diff,… → Design scenario: Generate a scenario… → Run the script: Pipe the scenario to uv… → …
  • Detecting regressions
  • SKILL.md covers When to Use BAT, Workflow, Output Structure and Scenario Design Guidelines, plus 5 more sections
  • Calls uv

What it does

Bat Adhoc is an agent skill from homeassistant-ai/ha-mcp. Run bot acceptance tests to validate MCP tools work correctly from a real AI agent's perspective. Use when testing PRs, detecting regressions, or verifying tool changes end-to-end with Claude/Gemini CLIs.

Its SKILL.md is about 1.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Testing & QA, covering End-to-end testing and MCP servers. The repository describes itself as: The Unofficial and Awesome Home Assistant MCP Server. The licence is MIT.

When your agent uses it

  • Detecting regressions
  • Verifying tool changes end-to-end with Claude/Gemini CLIs

Example prompts

  • “/bat-adhoc”

Requirements

  • Python 3
  • Pre-approved tools (allowed-tools): Bash, Read, Write

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Analyze the change: Read the diff, identify which tools are affected
  2. Design scenario: Generate a scenario JSON with setup/test/teardown prompts
  3. Run the script: Pipe the scenario to uv run python tests/uat/run_uat.py
  4. Evaluate summary: Check all_passed per agent. If true, you're done.
  5. Dig deeper on failure: Read results_file for full output, stderr, raw JSON
  6. Regression check: If test fails, re-run with --branch master to compare

What it can do on your machine

Read from SKILL.md and the folder at commit b274a93. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Bash
    • Read
    • Write

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bat Adhoc loads about 1.4k tokens when it runs. Until then it costs about 54 tokens; SKILL.md has 506 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~54
When it runs · the whole SKILL.md, loaded when a task matches
~1.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Bash, Read, Write

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from homeassistant-ai/ha-mcp at commit b274a93, republished under its MIT licence (© homeassistant-ai). 506 words, ~1,382 tokens.

Download SKILL.mdSave it as .claude/skills/bat-adhoc/SKILL.md (or your agent's skills folder).
name
bat-adhoc
description
Run bot acceptance tests to validate MCP tools work correctly from a real AI agent's perspective. Use when testing PRs, detecting regressions, or verifying tool changes end-to-end with Claude/Gemini CLIs.
allowed-tools
Bash, Read, Write
disable-model-invocation
true
argument-hint
scenario-description or --help

BAT - Bot Acceptance Testing

Bot acceptance testing validates that MCP tools work correctly from a real AI agent's perspective. You design test scenarios dynamically, run them via tests/uat/run_uat.py, and evaluate results.

When to Use BAT

  • PR validation: Test that tool changes work correctly from an agent's perspective
  • Regression detection: Compare behavior between branches
  • Integration verification: Ensure MCP tools work end-to-end with real agent CLIs

Workflow

  1. Analyze the change: Read the diff, identify which tools are affected
  2. Design scenario: Generate a scenario JSON with setup/test/teardown prompts
  3. Run the script: Pipe the scenario to uv run python tests/uat/run_uat.py
  4. Evaluate summary: Check all_passed per agent. If true, you're done.
  5. Dig deeper on failure: Read results_file for full output, stderr, raw JSON
  6. Regression check: If test fails, re-run with --branch master to compare

Output Structure

The runner returns a concise summary to stdout (saves context when all passes):

json
{
  "results_file": "/tmp/bat_results_abc123.json",
  "agents": {
    "gemini": {
      "all_passed": true,
      "test": {
        "completed": true,
        "duration_ms": 8100,
        "exit_code": 0,
        "num_turns": 5,
        "tool_stats": { "totalCalls": 4, "totalSuccess": 4, "totalFail": 0 }
      },
      "aggregate": {
        "total_duration_ms": 15300,
        "total_turns": 12,
        "total_tool_calls": 9,
        "total_tool_success": 9,
        "total_tool_fail": 0
      }
    }
  }
}
  • Phase stats: num_turns, tool_stats (per phase) for fine-grained comparison
  • Aggregate stats: Total counts across all phases for overall efficiency comparison
  • Output: every phase includes output (plus tool_trace when tool calls were logged); a failed phase also includes stderr when it is not empty
  • Full results: raw JSON, complete output always available at results_file

Scenario Design Guidelines

  • setup_prompt: Create any entities/state the test needs
  • test_prompt: Exercise the tools being tested, ask the agent to report results clearly
  • teardown_prompt: Clean up created entities
  • Keep prompts focused - each scenario tests ONE behavior
  • Ask the agent to report: what succeeded, what failed, any unexpected behavior

Example: Testing Error Signaling

bash
cat <<'EOF' | uv run python tests/uat/run_uat.py --agents gemini
{
  "setup_prompt": "Create a test automation called 'bat_error_test' with a time trigger at 23:59 and action to turn on light.bed_light.",
  "test_prompt": "Try to get automation 'automation.nonexistent_xyz'. Report if the tool signaled an error or returned a normal response. Then get automation 'automation.bat_error_test' and report its structure.",
  "teardown_prompt": "Delete automation 'bat_error_test' if it exists."
}
EOF
Show full SKILL.md (251 more words)Show less

Regression Comparison Workflow

Run the same scenario twice from the branch checkout and compare stats. --branch master installs ha-mcp from master on GitHub; omitting --branch runs the local code, which also covers unpushed commits:

bash
# Baseline: master
uv run python tests/uat/run_uat.py --scenario-file <scenario.json> --branch master --agents gemini

# Target: local code
uv run python tests/uat/run_uat.py --scenario-file <scenario.json> --agents gemini

To compare a pushed branch that is not checked out, pass --branch <branch> for the target run too.

Compare these metrics:

Primary (decide pass/fail on these):

  • Task completion: Did both pass? Any new failures?
  • Accuracy: Check agent output quality - did it understand the task correctly?
  • Tool success rate: Compare aggregate.total_tool_calls vs total_tool_fail

Secondary (report but don't decide on these alone):

  • Tool call count: Compare aggregate.total_tool_calls, aggregate.total_turns — directional signal, not conclusive (agent exploration varies between runs)
  • Duration: Compare aggregate.total_duration_ms — noisy due to network, cache misses, server load. Only flag large (>2x) regressions.

Robustness tip: Ask the same task in different ways (variation testing) to check if results are consistent across phrasings.

Cost Awareness

Each scenario invocation costs API credits (one per agent per phase). Design scenarios efficiently:

  • Combine related checks in a single test_prompt when possible
  • Only use setup/teardown when the test needs specific state
  • Start with one agent, expand to both only when cross-agent comparison matters

Handling Arguments

When /bat-adhoc is invoked with arguments:

If arguments contain a scenario description, generate the JSON scenario and run it:

/bat-adhoc test automation create with sunrise trigger then modify to sunset

→ Generate appropriate scenario JSON and execute

If --help or no arguments, show this help text.

Otherwise, treat $ARGUMENTS as instructions for what to test and design+run the scenario accordingly.

Full Documentation

For complete CLI reference and output format, see tests/uat/README.md.

© homeassistant-ai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/bat-adhoc of homeassistant-ai/ha-mcp.

Open the folder on GitHubat commit b274a93

Compare with similar skills

Bat Adhoc next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bat Adhoc compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bat Adhoc this skillhomeassistant-ai/ha-mcp5k—~1.4kAutomated safety check: NotesMIT
Qwen Code E2E TestingQwenLM/qwen-code28k—~2.1kAutomated safety check: PassApache-2.0
Glance TestDebugBase/glance156—~827Automated safety check: PassMIT
Edt MCP Build TestDitriXNew/EDT-MCP296—~1.4kAutomated safety check: PassAGPL-3.0
Tauri Dev MCP Bridgejerrywu001/cc-sessions-viewer396—~1.4kAutomated safety check: WarnMIT
Edt MCP E2E TestingDitriXNew/EDT-MCP296—~1.3kAutomated safety check: PassAGPL-3.0

Similar skills

  • Qwen Code E2E Testing

    QwenLM/qwen-code

    Guides end-to-end testing of the Qwen Code CLI in headless mode with real model calls, MCP test servers and inspection of raw API traffic.

    28k GitHub stars~2.1k tokensUpdated today
    Testing & QAAuto-check passed
  • Glance Test

    DebugBase/glance

    Run E2E browser tests on any web application using Glance MCP.

    156 GitHub stars~827 tokensUpdated 5 mo ago
    Testing & QAAuto-check passed
  • Edt MCP Build Test

    DitriXNew/EDT-MCP

    How to build the EDT-MCP Eclipse plugin (Tycho/Maven) and run its unit and e2e tests, plus the test conventions for this repo.

    296 GitHub stars~1.4k tokensUpdated 3 days ago
    Testing & QAAuto-check passed
  • Tauri Dev MCP Bridge

    jerrywu001/cc-sessions-viewer

    Launches a Tauri app's dev build with the MCP bridge compiled in, so an agent can take screenshots, read the live DOM, click real buttons and call real backend commands.

    396 GitHub stars~1.4k tokensUpdated 4 days ago
    Testing & QAAuto-check: warnings
  • Edt MCP E2E Testing

    DitriXNew/EDT-MCP

    How to write/run the AUTOMATED black-box e2e suite (tests/e2e/) that covers every EDT-MCP tool (62 today) against a live server with git-fixture isolation, happy + negative + error-quality coverage…

    296 GitHub stars~1.3k tokensUpdated 3 days ago
    Testing & QAAuto-check passed
  • Exploring The Wizard

    PostHog/wizard

    Official

    Drive the PostHog wizard headlessly against a throwaway app through wizard-ci MCP tools, inspect decisions, and capture the real TUI.

    197 GitHub stars~2.2k tokensUpdated yesterday
    Testing & QAAuto-check passed

More from homeassistant-ai/ha-mcp

All 8 skills in this repo
  • Contrib PR Review

    homeassistant-ai/ha-mcp

    Review a contribution PR for safety, quality, and readiness.

    5k GitHub stars~3.1k tokensUpdated today
    Auto-check: notes
  • Issue Analysis

    homeassistant-ai/ha-mcp

    Deep analysis of a single GitHub issue with codebase exploration, implementation planning, and architectural assessment.

    5k GitHub stars~753 tokensUpdated today
    Auto-check: notes
  • Issue To PR Resolver

    homeassistant-ai/ha-mcp

    Implement a GitHub issue end-to-end — create a worktree branch, implement the feature with tests, create a draft PR, then iteratively resolve all CI failures and review comments until the PR is clean.

    5k GitHub stars~1.2k tokensUpdated today
    Auto-check: notes
  • My PR Checker

    homeassistant-ai/ha-mcp

    Manage your own GitHub pull requests — check CI status, inline review comments, PR-level comments, resolve review threads, fix issues, and iterate until all checks pass and threads are resolved.

    5k GitHub stars~1.2k tokensUpdated today
    Auto-check: notes
  • Bat Story Eval

    homeassistant-ai/ha-mcp

    Compare MCP tool behavior between target and baseline versions using pre-built and custom stories with diff-based triage.

    5k GitHub stars~3.4k tokensUpdated today
    Auto-check: notes
  • Contributors Update

    homeassistant-ai/ha-mcp

    Find merged PR authors missing from README and update the contributors list after approval

    5k GitHub stars~983 tokensUpdated today
    Auto-check passed

Categories

Questions about Bat Adhoc

What does Bat Adhoc do?

Run bot acceptance tests to validate MCP tools work correctly from a real AI agent's perspective. Bat Adhoc is an agent skill from homeassistant-ai/ha-mcp. Run bot acceptance tests to validate MCP tools work correctly from a real AI agent's perspective.

When should I use Bat Adhoc?

Bat Adhoc fits situations like: detecting regressions; verifying tool changes end-to-end with Claude/Gemini CLIs.

How do I install Bat Adhoc in Claude Code?

Run `npx skills add homeassistant-ai/ha-mcp --skill bat-adhoc -a claude-code`. Or copy the skill folder (.claude/skills/bat-adhoc in homeassistant-ai/ha-mcp) into .claude/skills/bat-adhoc in your project. Claude Code loads it when a task matches its description.

How do I install Bat Adhoc in Codex?

Run `npx skills add homeassistant-ai/ha-mcp --skill bat-adhoc -a codex`. Or copy the skill folder (.claude/skills/bat-adhoc in homeassistant-ai/ha-mcp) into .agents/skills/bat-adhoc in your project. Codex loads it when a task matches its description.

Can I use Bat Adhoc in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add homeassistant-ai/ha-mcp --skill bat-adhoc -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bat-adhoc, .gemini/skills/bat-adhoc, .github/skills/bat-adhoc and .opencode/skills/bat-adhoc in your project.

What does Bat Adhoc need to run?

Going by SKILL.md and its folder, Bat Adhoc needs the command-line tools its instructions call (uv). Our summary lists: Python 3. Its frontmatter pre-approves these tools: Bash, Read, Write.

Does Bat Adhoc access the network?

SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Bat Adhoc safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Bat Adhoc use?

Bat Adhoc is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bat Adhoc use?

About 1.4k tokens (SKILL.md is roughly 5.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bat Adhoc?

Skills that share tags, products or a category with Bat Adhoc: Qwen Code E2E Testing (QwenLM/qwen-code, 28k stars), Glance Test (DebugBase/glance, 156 stars), Edt MCP Build Test (DitriXNew/EDT-MCP, 296 stars) and Tauri Dev MCP Bridge (jerrywu001/cc-sessions-viewer, 396 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bat Adhoc?

homeassistant-ai (a GitHub organization) maintains it in homeassistant-ai/ha-mcp, which has 5,015 GitHub stars. The repository holds 8 skills in this directory. The repository was last updated on October 11, 2026.

Source: homeassistant-ai/ha-mcp on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.