Official agent skill

Eval Authoring Skill

by Azure in Azure/azure-sdk-tools

Author and validate Vally evals for Agent Skills under .github/skills.

OfficialMITAuto-check passedAgent Workflows

Install Eval Authoring Skill

skills CLI
$ npx skills add Azure/azure-sdk-tools --skill eval-authoring-skill -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Azure/azure-sdk-tools eval-authoring-skill --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Azure/azure-sdk-tools.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.github/skills/eval-authoring-skill .claude/skills/eval-authoring-skill && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
eval-authoring-skill
GitHub stars
134
Token cost
~946 tokens
SKILL.md length
410 words
Files
2
Skills in repo
35
Repo updated
First seen
Licence
MIT

At a glance

Author and validate Vally evals for Agent Skills under .github/skills.

  • Works in 5 steps: Read the eval authoring guide… → Read the target SKILL.md — its WHEN/DO… → Write eval.yaml with routing stimuli… → …
  • Test skill routing
  • SKILL.md covers Triggers, Steps, Rules and References
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Eval Authoring Skill is an agent skill from Azure/azure-sdk-tools, published by the product's own GitHub organization. Author and validate Vally evals for Agent Skills under .github/skills. WHEN: "write a skill eval", "add trigger eval", "test skill routing", "add anti-trigger tests", "create skill capability eval", "harden skill graders". DO NOT USE FOR: single MCP tool prompt-to-tool coverage (use eval-authoring-tool), multi-tool or multi-turn scenarios (use eval-authoring-workflow), authoring SKILL.md itself (use skill-authoring).

Its SKILL.md is about 950 tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `evals/eval.yaml`). Compatibility notes: copilot-chat, @microsoft/vally-cli 0.14.0

It sits in Agent Workflows, covering Skill authoring, LLM evaluation and MCP servers. It works with GitHub. The repository describes itself as: Tools repository leveraged by the Azure SDK team. The licence is MIT.

When your agent uses it

  • Test skill routing
  • Add anti-trigger tests
  • Create skill capability eval
  • Harden skill graders

Example prompts

  • “write a skill eval”
  • “add trigger eval”
  • “test skill routing”
  • “/eval-authoring-skill”

Requirements

  • Compatibility (from SKILL.md): copilot-chat, @microsoft/vally-cli 0.14.0

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Read the eval authoring guide (.github/skills/eval-authoring/README.md) Step 0 to find this repo's vallyRoot/evalGlobs for the skill tier…
  2. Read the target SKILL.md — its WHEN/DO NOT USE FOR boundaries, invoked tools — and its existing evals/.
  3. Write eval.yaml with routing stimuli (trigger + anti-trigger, skill-invocation grader) and capability stimuli together in the same file…
  4. Supply concrete identifiers (repo path, package name, anything needed to reach the call) in every stimulus prompt — a vague prompt fails…
  5. Run strict eval-spec lint, then validate locally per the guide's "Running evals locally" section. Do not finish or open a PR until both…

What it can do on your machine

Read from SKILL.md and the folder at commit 6e4fb2b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    copilot-chat, @microsoft/vally-cli 0.14.0

    From compatibility in the SKILL.md frontmatter.

Context cost

Eval Authoring Skill loads about 946 tokens when it runs. Until then it costs about 110 tokens; SKILL.md has 410 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~110
When it runs · the whole SKILL.md, loaded when a task matches
~946

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Azure/azure-sdk-tools at commit 6e4fb2b, republished under its MIT licence (© Azure). 410 words, ~946 tokens.

Download SKILL.mdSave it as .claude/skills/eval-authoring-skill/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
eval-authoring-skill
description
Author and validate Vally evals for Agent Skills under .github/skills. WHEN: "write a skill eval", "add trigger eval", "test skill routing", "add anti-trigger tests", "create skill capability eval", "harden skill graders". DO NOT USE FOR: single MCP tool prompt-to-tool coverage (use eval-authoring-tool), multi-tool or multi-turn scenarios (use eval-authoring-workflow), authoring SKILL.md itself (use skill-authoring).
compatibility
copilot-chat, @microsoft/vally-cli 0.14.0
license
MIT
metadata.author
Microsoft
metadata.version
1.1.0

Skill Eval Authoring

Author Vally evals that verify one Agent Skill's routing and behavior in isolation, under .github/skills/<skill>/evals/. All the actual guidance — naming convention, placement, per-category requirements, glossary, grading patterns, grader catalog, anti-patterns, and worked examples — lives in one shared place so it never drifts across the three eval-authoring skills: the repository-local eval authoring guide at .github/skills/eval-authoring/README.md. This skill exists to route you there with skill-eval context already loaded, not to duplicate it.

Triggers

USE FOR: write a skill eval, add trigger eval, test skill routing, add anti-trigger tests, create skill capability eval, harden skill graders WHEN: "write a skill eval", "add trigger eval", "test skill routing", "add anti-trigger tests", "create skill capability eval", "harden skill graders" DO NOT USE FOR: single MCP tool prompt-to-tool coverage (use eval-authoring-tool), multi-tool or multi-turn scenarios (use eval-authoring-workflow), authoring SKILL.md itself (use skill-authoring)

Steps

  1. Read the eval authoring guide (.github/skills/eval-authoring/README.md) Step 0 to find this repo's vallyRoot/evalGlobs for the skill tier, then the guide's "Skill" column throughout (naming, requirements, worked example).
  2. Read the target SKILL.md — its WHEN/DO NOT USE FOR boundaries, invoked tools — and its existing evals/.
  3. Write eval.yaml with routing stimuli (trigger + anti-trigger, skill-invocation grader) and capability stimuli together in the same file per the guide's naming convention and four-layer pattern. Split into additional <behavior>.eval.yaml files only once a skill's coverage genuinely grows large. For a boundary prompt, mount and require the intended competing skill while disallowing the skill under test — an anti-trigger with no competing skill mounted trivially "passes" (guide anti-pattern A7).
  4. Supply concrete identifiers (repo path, package name, anything needed to reach the call) in every stimulus prompt — a vague prompt fails tool-calls with zero recorded calls, which looks like a routing bug but is really a missing-context prompt bug.
  5. Run strict eval-spec lint, then validate locally per the guide's "Running evals locally" section. Do not finish or open a PR until both lint and the focused local eval pass; inspect eval-results.md and results.jsonl on failure.
Show full SKILL.md (80 more words)Show less

Rules

  • Cover both activation and non-activation; anti-triggers should target neighboring skills, not random unrelated prompts.
  • Omit turn for conversation-wide assertions; pin a turn only when that turn owns the behavior unambiguously.
  • Never copy secrets or production-write identifiers into fixtures.
  • Update existing specs additively unless the requested contract changed.

References

  • Eval authoring guide: .github/skills/eval-authoring/README.md — naming, placement, requirements, glossary, graders, anti-patterns, worked examples, local commands. Shared with eval-authoring-tool/eval-authoring-workflow — update it there, not per-skill.
  • Repository-local eval README/configuration discovered in Step 0 (if present)

© Azure, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in .github/skills/eval-authoring-skill of Azure/azure-sdk-tools.

  • SKILL.md
  • evals/eval.yaml

Open the folder on GitHubat commit 6e4fb2b

Compare with similar skills

Eval Authoring Skill next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Eval Authoring Skill compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Eval Authoring Skill this skillAzure/azure-sdk-tools134—~946Automated safety check: PassMIT
Skill Seekers Builderyusufkaraaslan/Skill_Seekers15k—~760Automated safety check: PassMIT
Skill Creatorcuriositech/some_claude_skills243—~7.2kAutomated safety check: PassApache-2.0
Managed Deep Agentslangchain-ai/langchain-skills1.3k—~8.7kAutomated safety check: NotesMIT
Skill CreatorAzure/azqr79589 repos~8.2kAutomated safety check: PassApache-2.0
Microsoft Skill CreatorMicrosoftDocs/mcp1.9k3 repos~2.1kAutomated safety check: PassCC-BY-4.0

Similar skills

  • Skill Seekers Builder

    yusufkaraaslan/Skill_Seekers

    Detects the type of a knowledge source and uses the Skill Seekers MCP tools to turn docs, repos, PDFs or videos into packaged AI skills.

    15k GitHub stars~760 tokensUpdated 7 days ago
    Agent WorkflowsAuto-check passed
  • Skill Creator

    curiositech/some_claude_skills

    A skill your agent uses when creating a new Claude skill from scratch, editing or improving an existing skill, or measuring skill performance with evals and benchmarks.

    243 GitHub stars~7.2k tokensUpdated 1 mo ago
    Agent WorkflowsAuto-check passed
  • Managed Deep Agents

    langchain-ai/langchain-skills

    Official

    INVOKE THIS SKILL when building, testing, or deploying Managed Deep Agents in LangSmith.

    1.3k GitHub stars~8.7k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check: notes
  • Skill Creator

    Azure/azqr

    Official

    Create new skills, modify and improve existing skills, and measure skill performance.

    795 GitHub starsUsed in 89 repos~8.2k tokens
    Agent WorkflowsAuto-check passed
  • Microsoft Skill Creator

    MicrosoftDocs/mcp

    Official

    Create agent skills for Microsoft technologies using official documentation.

    1.9k GitHub starsUsed in 3 repos~2.1k tokens
    Agent WorkflowsAuto-check passed
  • Auto Skill Builder

    tradecatlabs/vibe-coding-cn

    Meta-skill that turns docs, APIs, code or specs into a reusable skill with references and a quality gate, and refactors skills that are unclear or misfire.

    17k GitHub starsUsed in 1 repo~2.4k tokens
    Agent WorkflowsAuto-check passed

More from Azure/azure-sdk-tools

All 35 skills in this repo
  • Apiview Feedback Resolution

    Azure/azure-sdk-tools

    Official

    Analyze and resolve APIView review feedback on Azure SDK PRs.

    134 GitHub stars~547 tokensUpdated today
    Auto-check passed
  • Official

    Deploy test resources and run Azure SDK tests in live, record, or playback mode.

    134 GitHub stars~1.5k tokensUpdated today
    Auto-check: notes
  • Azsdk Common Pipeline Analysis

    Azure/azure-sdk-tools

    Official

    Analyze Azure SDK CI/CD pipeline failures into a structured diagnosis, and define the required output format.

    134 GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Official

    Create, get, update, abandon, and link SDK PRs to release plan work items for Azure SDK releases.

    134 GitHub stars~2.7k tokensUpdated today
    Auto-check passed
  • Azure Typespec Assessment

    Azure/azure-sdk-tools

    Official

    Assess Azure TypeSpec Git diffs for semantic intent, REST and downstream SDK breaking changes, Azure Guidelines compliance, and documentation completeness.

    134 GitHub stars~1.9k tokensUpdated today
    Auto-check passed

Works with

Categories

Questions about Eval Authoring Skill

What does Eval Authoring Skill do?

Author and validate Vally evals for Agent Skills under .github/skills. Eval Authoring Skill is an agent skill from Azure/azure-sdk-tools, published by the product's own GitHub organization.github/skills.

When should I use Eval Authoring Skill?

Eval Authoring Skill fits situations like: test skill routing; add anti-trigger tests; create skill capability eval; harden skill graders.

How do I install Eval Authoring Skill in Claude Code?

Run `npx skills add Azure/azure-sdk-tools --skill eval-authoring-skill -a claude-code`. Or copy the skill folder (.github/skills/eval-authoring-skill in Azure/azure-sdk-tools) into .claude/skills/eval-authoring-skill in your project. Claude Code loads it when a task matches its description.

How do I install Eval Authoring Skill in Codex?

Run `npx skills add Azure/azure-sdk-tools --skill eval-authoring-skill -a codex`. Or copy the skill folder (.github/skills/eval-authoring-skill in Azure/azure-sdk-tools) into .agents/skills/eval-authoring-skill in your project. Codex loads it when a task matches its description.

Can I use Eval Authoring Skill in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Azure/azure-sdk-tools --skill eval-authoring-skill -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/eval-authoring-skill, .gemini/skills/eval-authoring-skill, .github/skills/eval-authoring-skill and .opencode/skills/eval-authoring-skill in your project.

What does Eval Authoring Skill need to run?

SKILL.md names no scripts, command-line tools or credentials: Eval Authoring Skill is instructions for the agent only. Compatibility (from SKILL.md): copilot-chat, @microsoft/vally-cli 0.14.0.

Does Eval Authoring Skill access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Eval Authoring Skill safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Eval Authoring Skill use?

Eval Authoring Skill is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Eval Authoring Skill use?

About 946 tokens (SKILL.md is roughly 3.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Eval Authoring Skill?

Skills that share tags, products or a category with Eval Authoring Skill: Skill Seekers Builder (yusufkaraaslan/Skill_Seekers, 15k stars), Skill Creator (curiositech/some_claude_skills, 243 stars), Managed Deep Agents (langchain-ai/langchain-skills, 1.3k stars) and Skill Creator (Azure/azqr, 795 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Eval Authoring Skill?

Azure (a GitHub organization, an official publisher) maintains it in Azure/azure-sdk-tools, which has 134 GitHub stars. The repository holds 35 skills in this directory. The repository was last updated on October 8, 2026.

Source: Azure/azure-sdk-tools on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.