Agent skill

Run Eval

by hidai25 in hidai25/eval-view

Run EvalView regression checks against golden baselines to detect regressions in AI agent behavior after code, prompt, or model changes.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Run Eval

skills CLI
$ npx skills add hidai25/eval-view --skill run-eval -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install hidai25/eval-view run-eval --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/hidai25/eval-view.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/run-eval .claude/skills/run-eval && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
run-eval
GitHub stars
137
Token cost
~557 tokens
SKILL.md length
272 words
Files
1
Skills in repo
4
Repo updated
First seen
Licence
Apache-2.0

At a glance

Run EvalView regression checks against golden baselines to detect regressions in AI agent behavior after code, prompt, or model changes.

  • Works in 5 steps: Locate the test directory. Look for… → Run a regression check using the… → Interpret results → …
  • Tasks that involve Building AI agents
  • SKILL.md covers What this does, Steps, CLI equivalent and Tips
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Run Eval is an agent skill from hidai25/eval-view. Run EvalView regression checks against golden baselines to detect regressions in AI agent behavior after code, prompt, or model changes.

Its SKILL.md is about 560 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering Building AI agents. It works with Python. The repository describes itself as: Regression testing for AI agents. Snapshot behavior,diff tool calls,catch regressions in CI. Works with LangGraph, CrewAI, OpenAI, Anthropic. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Building AI agents

Example prompts

  • “/run-eval”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Locate the test directory. Look for tests/evalview/ in the project. If it exists, use that. Otherwise check for a tests/ directory with…
  2. Run a regression check using the run_check MCP tool
  3. Interpret results
  4. If changes are intentional, offer to update the baseline by calling run_snapshot with an explanatory notes parameter.
  5. Generate a visual report (optional) by calling generate_visual_report for a detailed HTML breakdown of traces, diffs, scores, and timelines.

What it can do on your machine

Read from SKILL.md and the folder at commit 394f7d7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Run Eval loads about 557 tokens when it runs. Until then it costs about 36 tokens; SKILL.md has 272 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~36
When it runs · the whole SKILL.md, loaded when a task matches
~557

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from hidai25/eval-view at commit 394f7d7, republished under its Apache-2.0 licence (© hidai25). 272 words, ~557 tokens.

Download SKILL.mdSave it as .claude/skills/run-eval/SKILL.md (or your agent's skills folder).
name
run-eval
description
Run EvalView regression checks against golden baselines to detect regressions in AI agent behavior after code, prompt, or model changes.

Run Eval

Use this skill after making changes to an AI agent (prompt edits, model swaps, tool changes, code refactors) to verify nothing broke.

What this does

EvalView compares current agent behavior against saved golden baselines. It runs your test cases, evaluates the outputs, and reports a diff status for each test:

  • PASSED — behavior matches the baseline
  • OUTPUT_CHANGED — output shifted but may be intentional
  • TOOLS_CHANGED — different tools were called
  • REGRESSION — score dropped significantly (blocking failure)

Steps

  1. Locate the test directory. Look for tests/evalview/ in the project. If it exists, use that. Otherwise check for a tests/ directory with .yaml test files.

  2. Run a regression check using the run_check MCP tool:

    • If checking all tests: call run_check with the detected test_path
    • If checking a specific test: also pass the test parameter with the test name
  3. Interpret results:

    • If all tests pass, confirm to the user that no regressions were found
    • If REGRESSION is reported, show the diff (score delta, tool changes, output similarity) and offer to help fix it
    • If OUTPUT_CHANGED or TOOLS_CHANGED, flag it as a warning — the user should decide if the change is intentional
  4. If changes are intentional, offer to update the baseline by calling run_snapshot with an explanatory notes parameter.

  5. Generate a visual report (optional) by calling generate_visual_report for a detailed HTML breakdown of traces, diffs, scores, and timelines.

CLI equivalent

evalview check tests/evalview/
evalview check tests/evalview/ --test "my-test"
evalview snapshot tests/evalview/ --notes "updated after prompt refactor"

Tips

  • Use run_check frequently — it calls the Python API directly with no subprocess overhead.
  • A score delta near zero with TOOLS_CHANGED often means the agent found an equivalent path.
  • Always snapshot after confirming intentional changes so future checks compare against the new baseline.

© hidai25, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/run-eval of hidai25/eval-view.

Open the folder on GitHubat commit 394f7d7

Compare with similar skills

Run Eval next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Run Eval compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Run Eval this skillhidai25/eval-view137—~557Automated safety check: PassApache-2.0
Azure AI Projects Python SDKmicrosoft/skills3.1k6 repos~2.8kAutomated safety check: PassMIT
Senior Prompt Engineermaslennikov-ig/claude-code-orchestrator-kit2604 repos~1.4kAutomated safety check: PassCustom licence
Google Agents CLI Adk Codepifferologo/cloud-agents-cli1291 repos~768Automated safety check: PassApache-2.0
DSPy Language Model ProgrammingOrchestra-Research/AI-Research-SKILLs13k10 repos~3.8kAutomated safety check: PassMIT
E2b Code Interpreteragent-sandbox/agent-sandbox218—~2.3kAutomated safety check: PassApache-2.0

Similar skills

  • Official

    Reference for building on Microsoft Foundry with the azure-ai-projects Python SDK: project clients, versioned agents, evaluations, connections, datasets and indexes.

    3.1k GitHub starsUsed in 6 repos~2.8k tokens
    AI & LLM EngineeringAuto-check passed
  • Senior Prompt Engineer

    maslennikov-ig/claude-code-orchestrator-kit

    Provides reference guides and Python scripts for prompt optimization, RAG evaluation, and agent orchestration when building or tuning LLM systems.

    260 GitHub starsUsed in 4 repos~1.4k tokens
    AI & LLM EngineeringAuto-check passed
  • Google Agents CLI Adk Code

    pifferologo/cloud-agents-cli

    This skill should be used when the user wants to "write agent code", "build an agent with ADK", "add a tool", "create a callback", "define an agent", "use state management", or needs ADK (Agent…

    129 GitHub starsUsed in 1 repo~768 tokens
    AI & LLM EngineeringAuto-check passed
  • DSPy Language Model Programming

    Orchestra-Research/AI-Research-SKILLs

    Teaches an agent to build LM pipelines, RAG systems and agents in DSPy using signatures, modules and optimizers instead of hand-tuned prompts.

    13k GitHub starsUsed in 10 repos~3.8k tokens
    AI & LLM EngineeringAuto-check passed
  • E2b Code Interpreter

    agent-sandbox/agent-sandbox

    Execute code in E2B sandboxes and integrate with LLMs for tool calling.

    218 GitHub stars~2.3k tokensUpdated 18 days ago
    AI & LLM EngineeringAuto-check passed
  • Code Snippets

    MicrosoftDocs/semantic-kernel-docs

    Official

    How to reference code from sample repos in Agent Framework docs pages using :::code directives, snippet tags, zone pivots, and highlight attributes.

    264 GitHub stars~1.4k tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from hidai25/eval-view

  • Generate Tests

    hidai25/eval-view

    Generate EvalView test cases — either from a SKILL.md file using LLM-powered generation, or by capturing real agent interactions through a proxy.

    137 GitHub stars~658 tokensUpdated 1 mo ago
    Auto-check passed
  • Procrastination Buster

    hidai25/eval-view

    Beat procrastination with task breakdown, 2-minute starts, and accountability tracking

    137 GitHub starsUsed in 1 repo~837 tokens
    Auto-check passed
  • Watch

    hidai25/eval-view

    Start EvalView watch mode to automatically re-run regression checks whenever project files change.

    137 GitHub stars~577 tokensUpdated 1 mo ago
    Auto-check passed

Works with

Questions about Run Eval

What does Run Eval do?

Run EvalView regression checks against golden baselines to detect regressions in AI agent behavior after code, prompt, or model changes. Run Eval is an agent skill from hidai25/eval-view. Run EvalView regression checks against golden baselines to detect regressions in AI agent behavior after code, prompt, or model changes.

When should I use Run Eval?

Run Eval fits situations like: tasks that involve Building AI agents.

How do I install Run Eval in Claude Code?

Run `npx skills add hidai25/eval-view --skill run-eval -a claude-code`. Or copy the skill folder (skills/run-eval in hidai25/eval-view) into .claude/skills/run-eval in your project. Claude Code loads it when a task matches its description.

How do I install Run Eval in Codex?

Run `npx skills add hidai25/eval-view --skill run-eval -a codex`. Or copy the skill folder (skills/run-eval in hidai25/eval-view) into .agents/skills/run-eval in your project. Codex loads it when a task matches its description.

Can I use Run Eval in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add hidai25/eval-view --skill run-eval -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/run-eval, .gemini/skills/run-eval, .github/skills/run-eval and .opencode/skills/run-eval in your project.

What does Run Eval need to run?

SKILL.md names no scripts, command-line tools or credentials: Run Eval is instructions for the agent only. Our summary lists: Python 3.

Does Run Eval access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Run Eval safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Run Eval use?

Run Eval is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Run Eval use?

About 557 tokens (SKILL.md is roughly 2.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Run Eval?

Skills that share tags, products or a category with Run Eval: Azure AI Projects Python SDK (microsoft/skills, 3.1k stars), Senior Prompt Engineer (maslennikov-ig/claude-code-orchestrator-kit, 260 stars), Google Agents CLI Adk Code (pifferologo/cloud-agents-cli, 129 stars) and DSPy Language Model Programming (Orchestra-Research/AI-Research-SKILLs, 13k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Run Eval?

hidai25 (a GitHub user) maintains it in hidai25/eval-view, which has 137 GitHub stars. The repository holds 4 skills in this directory. The repository was last updated on September 5, 2026.

Source: hidai25/eval-view on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.