Agent skill

Generate Tests

by hidai25 in hidai25/eval-view

Generate EvalView test cases — either from a SKILL.md file using LLM-powered generation, or by capturing real agent interactions through a proxy.

Apache-2.0Auto-check passedTesting & QA

Install Generate Tests

skills CLI
$ npx skills add hidai25/eval-view --skill generate-tests -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install hidai25/eval-view generate-tests --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/hidai25/eval-view.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/generate-tests .claude/skills/generate-tests && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
generate-tests
GitHub stars
137
Token cost
~658 tokens
SKILL.md length
270 words
Files
1
Skills in repo
4
Repo updated
First seen
Licence
Apache-2.0

At a glance

Generate EvalView test cases — either from a SKILL.md file using LLM-powered generation, or by capturing real agent interactions through a proxy.

  • Works in 4 steps: Generate tests from a SKILL.md file → Create individual test cases manually → Capture real interactions → …
  • Tasks that involve Test generation
  • SKILL.md covers Four approaches and Running generated tests
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Generate Tests is an agent skill from hidai25/eval-view. Generate EvalView test cases — either from a SKILL.md file using LLM-powered generation, or by capturing real agent interactions through a proxy.

Its SKILL.md is about 660 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Testing & QA, covering Test generation, Skill authoring and Building AI agents. The repository describes itself as: Regression testing for AI agents. Snapshot behavior,diff tool calls,catch regressions in CI. Works with LangGraph, CrewAI, OpenAI, Anthropic. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Test generation
  • Tasks that involve Skill authoring
  • Tasks that involve Building AI agents

Example prompts

  • “/generate-tests”

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Generate tests from a SKILL.md file
  2. Create individual test cases manually
  3. Capture real interactions
  4. Validate a skill before testing

What it can do on your machine

Read from SKILL.md and the folder at commit 394f7d7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Generate Tests loads about 658 tokens when it runs. Until then it costs about 40 tokens; SKILL.md has 270 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~40
When it runs · the whole SKILL.md, loaded when a task matches
~658

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from hidai25/eval-view at commit 394f7d7, republished under its Apache-2.0 licence (© hidai25). 270 words, ~658 tokens.

Download SKILL.mdSave it as .claude/skills/generate-tests/SKILL.md (or your agent's skills folder).
name
generate-tests
description
Generate EvalView test cases — either from a SKILL.md file using LLM-powered generation, or by capturing real agent interactions through a proxy.

Generate Tests

Use this skill when the user wants to create test cases for their AI agent or skill without writing YAML by hand.

Four approaches

1. Generate tests from a SKILL.md file

Use the generate_skill_tests MCP tool to auto-generate a test suite from a skill definition. This reads the SKILL.md and produces YAML test cases covering explicit triggers, implicit triggers, contextual triggers, and negative cases.

Steps:

  1. Ask the user which SKILL.md to generate tests for (or detect it from context).
  2. Call generate_skill_tests with:
    • skill_path: path to the SKILL.md file
    • output_path (optional): where to save the generated YAML
    • count (optional): number of test cases (default: 10)
  3. After generation, offer to run the tests with run_skill_test.

CLI equivalent:

evalview skill generate-tests .claude/skills/my-skill/SKILL.md --auto
evalview skill generate-tests .claude/skills/my-skill/SKILL.md -c 20 -o tests/my-skill-tests.yaml
2. Create individual test cases manually

Use the create_test MCP tool to create a single test YAML file from a description.

Steps:

  1. Gather from the user: test name, query, expected tools, forbidden tools, expected output keywords, and minimum score.
  2. Call create_test with the parameters.
  3. After creating the test, call run_snapshot to establish the golden baseline.
3. Capture real interactions

Use the CLI evalview capture command to proxy real agent traffic and save interactions as test YAMLs automatically. This records the query, output, and tool calls from live usage.

CLI equivalent:

evalview capture --agent http://localhost:8080/execute --output-dir tests/test-cases
evalview capture --multi-turn  # saves all turns as one multi-turn conversation test
4. Validate a skill before testing

Use validate_skill to check a SKILL.md for correct structure and completeness before generating tests from it.

Running generated tests

After generating tests, execute them with run_skill_test:

  • test_file: path to the generated YAML
  • no_rubric: true for fast deterministic-only checks (no LLM cost)
  • verbose: true for detailed output on all tests

CLI equivalent:

evalview skill test tests/my-skill-tests.yaml
evalview skill test tests/my-skill-tests.yaml --no-rubric  # fast, $0
evalview skill test tests/my-skill-tests.yaml --verbose --model claude-sonnet-4-20250514

© hidai25, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/generate-tests of hidai25/eval-view.

Open the folder on GitHubat commit 394f7d7

Compare with similar skills

Generate Tests next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Generate Tests compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Generate Tests this skillhidai25/eval-view137—~658Automated safety check: PassApache-2.0
Langgraph Testing Evaluationsoba-labs/langchain-agent-skills107—~2.3kAutomated safety check: PassMIT
Create Agent Templateharness/harness-skills115—~2.2kAutomated safety check: PassApache-2.0
Agent Creatorjdforsythe/forge151—~4.5kAutomated safety check: PassMIT
Veomni Patchgen ModelByteDance-Seed/VeOmni2.2k—~9.6kAutomated safety check: PassApache-2.0
Guidancetestdouble/han279—~1.8kAutomated safety check: PassMIT

Similar skills

  • Langgraph Testing Evaluation

    soba-labs/langchain-agent-skills

    A skill your agent uses when you need to test or evaluate LangGraph/LangChain agents: writing unit or integration tests, generating test scaffolds, mocking LLM/tool behavior, running trajectory…

    107 GitHub stars~2.3k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Create Agent Template

    harness/harness-skills

    Generate Harness Agent Template files for AI-powered automation agents.

    115 GitHub stars~2.2k tokensUpdated yesterday
    DevelopmentAuto-check passed
  • Agent Creator

    jdforsythe/forge

    Creates structured agent definitions using the 7-component format grounded in persona science (the alignment-accuracy tradeoff), vocabulary routing, and the MAST failure taxonomy + Forge watchlist.

    151 GitHub stars~4.5k tokensUpdated 3 mo ago
    Agent WorkflowsAuto-check passed
  • Veomni Patchgen Model

    ByteDance-Seed/VeOmni

    Author or refresh a VeOmni model's patchgen-generated modeling under generated/ — GPU and/or NPU config, dense or MoE, text / VLM / Omni.

    2.2k GitHub stars~9.6k tokensUpdated today
    Testing & QAAuto-check passed
  • Guidance

    testdouble/han

    Authoritative guidance for building Claude Code skills, agents, and plugins, plus init and update steps that install and refresh the plugin-building skills in the current repository.

    279 GitHub stars~1.8k tokensUpdated 7 days ago
    Agent WorkflowsAuto-check passed
  • Eval Guide

    microsoft/eval-guide

    Official

    Eval enablement accelerator — help customers think through "what does good look like" for their AI agent, then generate a structured eval plan and test cases they can use immediately.

    138 GitHub stars~22k tokensUpdated 3 mo ago
    Testing & QAAuto-check: warnings

More from hidai25/eval-view

  • Run Eval

    hidai25/eval-view

    Run EvalView regression checks against golden baselines to detect regressions in AI agent behavior after code, prompt, or model changes.

    137 GitHub stars~557 tokensUpdated 1 mo ago
    Auto-check passed
  • Procrastination Buster

    hidai25/eval-view

    Beat procrastination with task breakdown, 2-minute starts, and accountability tracking

    137 GitHub starsUsed in 1 repo~837 tokens
    Auto-check passed
  • Watch

    hidai25/eval-view

    Start EvalView watch mode to automatically re-run regression checks whenever project files change.

    137 GitHub stars~577 tokensUpdated 1 mo ago
    Auto-check passed

Questions about Generate Tests

What does Generate Tests do?

Generate EvalView test cases — either from a SKILL.md file using LLM-powered generation, or by capturing real agent interactions through a proxy. Generate Tests is an agent skill from hidai25/eval-view.md file using LLM-powered generation, or by capturing real agent interactions through a proxy.

When should I use Generate Tests?

Generate Tests fits situations like: tasks that involve Test generation; tasks that involve Skill authoring; tasks that involve Building AI agents.

How do I install Generate Tests in Claude Code?

Run `npx skills add hidai25/eval-view --skill generate-tests -a claude-code`. Or copy the skill folder (skills/generate-tests in hidai25/eval-view) into .claude/skills/generate-tests in your project. Claude Code loads it when a task matches its description.

How do I install Generate Tests in Codex?

Run `npx skills add hidai25/eval-view --skill generate-tests -a codex`. Or copy the skill folder (skills/generate-tests in hidai25/eval-view) into .agents/skills/generate-tests in your project. Codex loads it when a task matches its description.

Can I use Generate Tests in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add hidai25/eval-view --skill generate-tests -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/generate-tests, .gemini/skills/generate-tests, .github/skills/generate-tests and .opencode/skills/generate-tests in your project.

What does Generate Tests need to run?

SKILL.md names no scripts, command-line tools or credentials: Generate Tests is instructions for the agent only.

Does Generate Tests access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Generate Tests safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Generate Tests use?

Generate Tests is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Generate Tests use?

About 658 tokens (SKILL.md is roughly 2.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Generate Tests?

Skills that share tags, products or a category with Generate Tests: Langgraph Testing Evaluation (soba-labs/langchain-agent-skills, 107 stars), Create Agent Template (harness/harness-skills, 115 stars), Agent Creator (jdforsythe/forge, 151 stars) and Veomni Patchgen Model (ByteDance-Seed/VeOmni, 2.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Generate Tests?

hidai25 (a GitHub user) maintains it in hidai25/eval-view, which has 137 GitHub stars. The repository holds 4 skills in this directory. The repository was last updated on September 5, 2026.

Source: hidai25/eval-view on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.