Official agent skill

Eval Authoring Tool

by Azure in Azure/azure-sdk-tools

Author and validate hermetic single-tool Vally evals under evals/tools.

OfficialMITAuto-check passedAI & LLM Engineering

Install Eval Authoring Tool

skills CLI
$ npx skills add Azure/azure-sdk-tools --skill eval-authoring-tool -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Azure/azure-sdk-tools eval-authoring-tool --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Azure/azure-sdk-tools.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.github/skills/eval-authoring-tool .claude/skills/eval-authoring-tool && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
eval-authoring-tool
GitHub stars
134
Token cost
~893 tokens
SKILL.md length
391 words
Files
2
Skills in repo
35
Repo updated
First seen
Licence
MIT

At a glance

Author and validate hermetic single-tool Vally evals under evals/tools.

  • Works in 6 steps: Read the eval authoring guide… → Confirm the scenario expects one primary… → Add stimuli to the matching… → …
  • : skill routing
  • SKILL.md covers Triggers, Steps, Rules and References
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Eval Authoring Tool is an agent skill from Azure/azure-sdk-tools, published by the product's own GitHub organization. Author and validate hermetic single-tool Vally evals under evals/tools. WHEN: "write a tool eval", "add prompt-to-tool coverage", "test MCP tool selection", "add tool catalog eval", "create single-tool scenario", "harden tool-call grader". DO NOT USE FOR: skill routing or capability evals (use eval-authoring-skill), multi-tool, multi-turn, or live scenarios (use eval-authoring-workflow).

Its SKILL.md is about 890 tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `evals/eval.yaml`). Compatibility notes: copilot-chat, @microsoft/vally-cli 0.14.0

It sits in AI & LLM Engineering, covering LLM evaluation and MCP servers. The repository describes itself as: Tools repository leveraged by the Azure SDK team. The licence is MIT.

When your agent uses it

  • : skill routing
  • Capability evals (use eval-authoring-skill)
  • Live scenarios (use eval-authoring-workflow)

Example prompts

  • “write a tool eval”
  • “add prompt-to-tool coverage”
  • “test MCP tool selection”
  • “/eval-authoring-tool”

Requirements

  • Compatibility (from SKILL.md): copilot-chat, @microsoft/vally-cli 0.14.0

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Read the eval authoring guide (.github/skills/eval-authoring/README.md) Step 0 to find this repo's vallyRoot/evalGlobs for the tool tier…
  2. Confirm the scenario expects one primary tool. If success requires orchestration, conversation state, or live services, use…
  3. Add stimuli to the matching prompt-to-tool-.eval.yaml; create a separate file only when it needs fixtures or outcome grading.
  4. Use realistic, concrete prompts with multiple natural phrasings and collision cases that disallow the nearest competing tool. Make…
  5. Keep the unit tier hermetic (environment: azsdk-mcp-mock, tags.tier: unit, established area tag). Avoid git worktrees and production writes.
  6. Run strict eval-spec lint, then validate locally per the guide's "Running evals locally" section, building the mock MCP first. Do not…

What it can do on your machine

Read from SKILL.md and the folder at commit 6e4fb2b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    copilot-chat, @microsoft/vally-cli 0.14.0

    From compatibility in the SKILL.md frontmatter.

Context cost

Eval Authoring Tool loads about 893 tokens when it runs. Until then it costs about 103 tokens; SKILL.md has 391 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~103
When it runs · the whole SKILL.md, loaded when a task matches
~893

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Azure/azure-sdk-tools at commit 6e4fb2b, republished under its MIT licence (© Azure). 391 words, ~893 tokens.

Download SKILL.mdSave it as .claude/skills/eval-authoring-tool/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
eval-authoring-tool
description
Author and validate hermetic single-tool Vally evals under evals/tools. WHEN: "write a tool eval", "add prompt-to-tool coverage", "test MCP tool selection", "add tool catalog eval", "create single-tool scenario", "harden tool-call grader". DO NOT USE FOR: skill routing or capability evals (use eval-authoring-skill), multi-tool, multi-turn, or live scenarios (use eval-authoring-workflow).
compatibility
copilot-chat, @microsoft/vally-cli 0.14.0
license
MIT
metadata.author
Microsoft
metadata.version
1.1.0

Tool Eval Authoring

Author Vally evals that verify one MCP tool is selected correctly for a given prompt, under evals/tools/. All the actual guidance — naming convention, placement, per-category requirements, glossary, grading patterns, grader catalog, anti-patterns, and worked examples — lives in one shared place so it never drifts across the three eval-authoring skills: the repository-local eval authoring guide at .github/skills/eval-authoring/README.md. This skill exists to route you there with tool-eval context already loaded, not to duplicate it.

Triggers

USE FOR: write a tool eval, add prompt-to-tool coverage, test MCP tool selection, add tool catalog eval, create single-tool scenario, harden tool-call grader WHEN: "write a tool eval", "add prompt-to-tool coverage", "test MCP tool selection", "add tool catalog eval", "create single-tool scenario", "harden tool-call grader" DO NOT USE FOR: skill routing or capability evals (use eval-authoring-skill), multi-tool, multi-turn, or live scenarios (use eval-authoring-workflow)

Steps

  1. Read the eval authoring guide (.github/skills/eval-authoring/README.md) Step 0 to find this repo's vallyRoot/evalGlobs for the tool tier, then the guide's "Tool" column throughout (naming, requirements, worked example).
  2. Confirm the scenario expects one primary tool. If success requires orchestration, conversation state, or live services, use eval-authoring-workflow instead.
  3. Add stimuli to the matching prompt-to-tool-<area>.eval.yaml; create a separate file only when it needs fixtures or outcome grading.
  4. Use realistic, concrete prompts with multiple natural phrasings and collision cases that disallow the nearest competing tool. Make tool-calls the primary signal; avoid the anti-patterns in the guide (vacuous keyword-only grading, missing scoring.threshold).
  5. Keep the unit tier hermetic (environment: azsdk-mcp-mock, tags.tier: unit, established area tag). Avoid git worktrees and production writes.
  6. Run strict eval-spec lint, then validate locally per the guide's "Running evals locally" section, building the mock MCP first. Do not finish or open a PR until both lint and the focused eval pass; inspect recorded tool names and arguments on failure.
Show full SKILL.md (88 more words)Show less

Rules

  • Use exact MCP tool names and assert forbidden alternatives where ambiguity exists.
  • Do not force a tool through unnatural prompt instructions unless direct invocation is the contract being tested.
  • Keep fixture paths relative to the eval file and fixture data minimal and non-secret.
  • Preserve existing namespace coverage and add new stimuli instead of duplicating files.

References

  • Eval authoring guide: .github/skills/eval-authoring/README.md — naming, placement, requirements, glossary, graders, anti-patterns, worked examples, local commands. Shared with eval-authoring-skill/eval-authoring-workflow — update it there, not per-skill.
  • Repository-local eval README/configuration discovered in Step 0 (if present)

© Azure, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in .github/skills/eval-authoring-tool of Azure/azure-sdk-tools.

  • SKILL.md
  • evals/eval.yaml

Open the folder on GitHubat commit 6e4fb2b

Compare with similar skills

Eval Authoring Tool next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Eval Authoring Tool compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Eval Authoring Tool this skillAzure/azure-sdk-tools134—~893Automated safety check: PassMIT
Opik Online Evalcomet-ml/opik-mcp219—~3kAutomated safety check: NotesApache-2.0
Mastramajiayu000/claude-skill-registry6661 repos~3.2kAutomated safety check: PassMIT
Agent Eval Casesagentailor/fullstack-langgraph-nextjs-agent132—~5.3kAutomated safety check: PassMIT
Improving MCP ToolsPostHog/posthog40k—~1.5kAutomated safety check: PassCustom licence
Managed Deep Agentslangchain-ai/langchain-skills1.3k—~8.7kAutomated safety check: NotesMIT

Similar skills

  • Opik Online Eval

    comet-ml/opik-mcp

    Take a judge live on production traffic — create an Opik online evaluation rule (LLM-as-judge or Python metric) on a project with sampling, filters, variable mapping, and a cost cap, then confirm…

    219 GitHub stars~3k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check: notes
  • Mastra

    majiayu000/claude-skill-registry

    A skill your agent uses when working with Mastra - the TypeScript AI framework for building agents, workflows, tools, and AI-powered applications.

    666 GitHub starsUsed in 1 repo~3.2k tokens
    AI & LLM EngineeringAuto-check passed
  • Agent Eval Cases

    agentailor/fullstack-langgraph-nextjs-agent

    Decide which AI agent behaviors are worth an eval case, then write those cases — harness-, framework-, and language-agnostic.

    132 GitHub stars~5.3k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Improving MCP Tools

    PostHog/posthog

    Official

    Run an improve-my-MCP campaign: an autoresearch-style loop that measures the MCP agent experience with the eval harness, picks the highest-impact tool problem from production data, makes one bounded…

    40k GitHub stars~1.5k tokensUpdated today
    Agent WorkflowsAuto-check passed
  • Managed Deep Agents

    langchain-ai/langchain-skills

    Official

    INVOKE THIS SKILL when building, testing, or deploying Managed Deep Agents in LangSmith.

    1.3k GitHub stars~8.7k tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check: notes
  • Conventions MCP

    stella/stella

    Apply when adding or changing an MCP tool, a capability the CLI generates, a tool input or output schema, a tool description, an error envelope, or an agent-facing reference resource.

    258 GitHub stars~2.7k tokensUpdated today
    Agent WorkflowsAuto-check passed

More from Azure/azure-sdk-tools

All 35 skills in this repo
  • Apiview Feedback Resolution

    Azure/azure-sdk-tools

    Official

    Analyze and resolve APIView review feedback on Azure SDK PRs.

    134 GitHub stars~547 tokensUpdated today
    Auto-check passed
  • Official

    Deploy test resources and run Azure SDK tests in live, record, or playback mode.

    134 GitHub stars~1.5k tokensUpdated today
    Auto-check: notes
  • Azsdk Common Pipeline Analysis

    Azure/azure-sdk-tools

    Official

    Analyze Azure SDK CI/CD pipeline failures into a structured diagnosis, and define the required output format.

    134 GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Official

    Create, get, update, abandon, and link SDK PRs to release plan work items for Azure SDK releases.

    134 GitHub stars~2.7k tokensUpdated today
    Auto-check passed
  • Azure Typespec Assessment

    Azure/azure-sdk-tools

    Official

    Assess Azure TypeSpec Git diffs for semantic intent, REST and downstream SDK breaking changes, Azure Guidelines compliance, and documentation completeness.

    134 GitHub stars~1.9k tokensUpdated today
    Auto-check passed

Questions about Eval Authoring Tool

What does Eval Authoring Tool do?

Author and validate hermetic single-tool Vally evals under evals/tools. Eval Authoring Tool is an agent skill from Azure/azure-sdk-tools, published by the product's own GitHub organization. Author and validate hermetic single-tool Vally evals under evals/tools.

When should I use Eval Authoring Tool?

Eval Authoring Tool fits situations like: : skill routing; capability evals (use eval-authoring-skill); live scenarios (use eval-authoring-workflow).

How do I install Eval Authoring Tool in Claude Code?

Run `npx skills add Azure/azure-sdk-tools --skill eval-authoring-tool -a claude-code`. Or copy the skill folder (.github/skills/eval-authoring-tool in Azure/azure-sdk-tools) into .claude/skills/eval-authoring-tool in your project. Claude Code loads it when a task matches its description.

How do I install Eval Authoring Tool in Codex?

Run `npx skills add Azure/azure-sdk-tools --skill eval-authoring-tool -a codex`. Or copy the skill folder (.github/skills/eval-authoring-tool in Azure/azure-sdk-tools) into .agents/skills/eval-authoring-tool in your project. Codex loads it when a task matches its description.

Can I use Eval Authoring Tool in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Azure/azure-sdk-tools --skill eval-authoring-tool -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/eval-authoring-tool, .gemini/skills/eval-authoring-tool, .github/skills/eval-authoring-tool and .opencode/skills/eval-authoring-tool in your project.

What does Eval Authoring Tool need to run?

SKILL.md names no scripts, command-line tools or credentials: Eval Authoring Tool is instructions for the agent only. Compatibility (from SKILL.md): copilot-chat, @microsoft/vally-cli 0.14.0.

Does Eval Authoring Tool access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Eval Authoring Tool safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Eval Authoring Tool use?

Eval Authoring Tool is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Eval Authoring Tool use?

About 893 tokens (SKILL.md is roughly 3.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Eval Authoring Tool?

Skills that share tags, products or a category with Eval Authoring Tool: Opik Online Eval (comet-ml/opik-mcp, 219 stars), Mastra (majiayu000/claude-skill-registry, 666 stars), Agent Eval Cases (agentailor/fullstack-langgraph-nextjs-agent, 132 stars) and Improving MCP Tools (PostHog/posthog, 40k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Eval Authoring Tool?

Azure (a GitHub organization, an official publisher) maintains it in Azure/azure-sdk-tools, which has 134 GitHub stars. The repository holds 35 skills in this directory. The repository was last updated on October 8, 2026.

Source: Azure/azure-sdk-tools on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.