Official agent skill

Eval Authoring Workflow

by Azure in Azure/azure-sdk-tools

Author and validate multi-tool, multi-turn, mock, and live Vally scenarios under evals/workflows.

OfficialMITAuto-check passedAI & LLM Engineering

Install Eval Authoring Workflow

skills CLI
$ npx skills add Azure/azure-sdk-tools --skill eval-authoring-workflow -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Azure/azure-sdk-tools eval-authoring-workflow --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Azure/azure-sdk-tools.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.github/skills/eval-authoring-workflow .claude/skills/eval-authoring-workflow && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
eval-authoring-workflow
GitHub stars
134
Token cost
~972 tokens
SKILL.md length
431 words
Files
2
Skills in repo
35
Repo updated
First seen
Licence
MIT

At a glance

Author and validate multi-tool, multi-turn, mock, and live Vally scenarios under evals/workflows.

  • Works in 6 steps: Read the eval authoring guide… → Use evals/workflows/mock/ by default.… → Model one user goal per stimulus. Use… → …
  • : per-skill routing/capability evals (use eval-authoring-skill)
  • SKILL.md covers Triggers, Steps, Rules and References
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Eval Authoring Workflow is an agent skill from Azure/azure-sdk-tools, published by the product's own GitHub organization. Author and validate multi-tool, multi-turn, mock, and live Vally scenarios under evals/workflows. WHEN: "write a workflow eval", "add end-to-end eval", "test tool sequence", "create multi-turn eval", "add mock workflow scenario", "add live eval". DO NOT USE FOR: per-skill routing/capability evals (use eval-authoring-skill), isolated prompt-to-tool tests (use eval-authoring-tool).

Its SKILL.md is about 970 tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `evals/eval.yaml`). Compatibility notes: copilot-chat, @microsoft/vally-cli 0.14.0

It sits in AI & LLM Engineering, covering LLM evaluation and End-to-end testing. The repository describes itself as: Tools repository leveraged by the Azure SDK team. The licence is MIT.

When your agent uses it

  • : per-skill routing/capability evals (use eval-authoring-skill)
  • Isolated prompt-to-tool tests (use eval-authoring-tool)

Example prompts

  • “write a workflow eval”
  • “add end-to-end eval”
  • “test tool sequence”
  • “/eval-authoring-workflow”

Requirements

  • Compatibility (from SKILL.md): copilot-chat, @microsoft/vally-cli 0.14.0

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Read the eval authoring guide (.github/skills/eval-authoring/README.md) Step 0 to find this repo's vallyRoot/evalGlobs for the workflow…
  2. Use evals/workflows/mock/ by default. Choose live/ only when the behavior cannot be represented by the mock MCP; document writes…
  3. Model one user goal per stimulus. Use turns only when conversation state matters; otherwise keep a single prompt. Mount every candidate…
  4. Combine process and outcome graders per the guide's four-layer pattern and grader catalog: required/disallowed skills and tools, ordering…
  5. Scope graders to turn only when that turn owns the assertion unambiguously; otherwise grade the full conversation.
  6. Materialize generated fixture sources, run strict eval-spec lint, then validate locally per the guide's "Running evals locally" section…

What it can do on your machine

Read from SKILL.md and the folder at commit 6e4fb2b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    copilot-chat, @microsoft/vally-cli 0.14.0

    From compatibility in the SKILL.md frontmatter.

Context cost

Eval Authoring Workflow loads about 972 tokens when it runs. Until then it costs about 102 tokens; SKILL.md has 431 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~102
When it runs · the whole SKILL.md, loaded when a task matches
~972

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Azure/azure-sdk-tools at commit 6e4fb2b, republished under its MIT licence (© Azure). 431 words, ~972 tokens.

Download SKILL.mdSave it as .claude/skills/eval-authoring-workflow/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
eval-authoring-workflow
description
Author and validate multi-tool, multi-turn, mock, and live Vally scenarios under evals/workflows. WHEN: "write a workflow eval", "add end-to-end eval", "test tool sequence", "create multi-turn eval", "add mock workflow scenario", "add live eval". DO NOT USE FOR: per-skill routing/capability evals (use eval-authoring-skill), isolated prompt-to-tool tests (use eval-authoring-tool).
compatibility
copilot-chat, @microsoft/vally-cli 0.14.0
license
MIT
metadata.author
Microsoft
metadata.version
1.1.0

Workflow Eval Authoring

Author Vally evals that verify multi-skill, multi-tool, or multi-turn orchestration, state, and outcomes across an entire agent conversation, under evals/workflows/. All the actual guidance — naming convention, placement, per-category requirements, glossary, grading patterns, grader catalog, anti-patterns, and worked examples — lives in one shared place so it never drifts across the three eval-authoring skills: the repository-local eval authoring guide at .github/skills/eval-authoring/README.md. This skill exists to route you there with workflow-eval context already loaded, not to duplicate it.

Triggers

USE FOR: write a workflow eval, add end-to-end eval, test tool sequence, create multi-turn eval, add mock workflow scenario, add live eval WHEN: "write a workflow eval", "add end-to-end eval", "test tool sequence", "create multi-turn eval", "add mock workflow scenario", "add live eval" DO NOT USE FOR: per-skill routing/capability evals (use eval-authoring-skill), isolated prompt-to-tool tests (use eval-authoring-tool)

Steps

  1. Read the eval authoring guide (.github/skills/eval-authoring/README.md) Step 0 to find this repo's vallyRoot/evalGlobs for the workflow tier, then the guide's "Workflow" column throughout (naming, requirements, worked example).
  2. Use evals/workflows/mock/ by default. Choose live/ only when the behavior cannot be represented by the mock MCP; document writes, authentication, cleanup, and nightly-only execution.
  3. Model one user goal per stimulus. Use turns only when conversation state matters; otherwise keep a single prompt. Mount every candidate skill explicitly, provide minimal file/git fixtures, and bound turns, tokens, workers, and timeout.
  4. Combine process and outcome graders per the guide's four-layer pattern and grader catalog: required/disallowed skills and tools, ordering where essential, files/commands, response quality. Avoid overfitting to incidental call sequences and the guide's anti-patterns (e.g. boundary anti-triggers with no competing skill mounted).
  5. Scope graders to turn only when that turn owns the assertion unambiguously; otherwise grade the full conversation.
  6. Materialize generated fixture sources, run strict eval-spec lint, then validate locally per the guide's "Running evals locally" section. Build the matching MCP and prime git fixtures when required. Do not finish or open a PR until both lint and the focused eval pass; inspect every turn and tool call on failure.
Show full SKILL.md (94 more words)Show less

Rules

  • Keep mock workflows hermetic and repeatable; use fake IDs and canned responses.
  • Never run a live write scenario without the documented test area and safety environment.
  • Preserve relative path invariants for environment.skills, files, and git.source; every environment.files entry needs both src and dest.
  • Weight by grader type, include every used type (use 0 only intentionally), and normalize positive weights to 1.0.

References

  • Eval authoring guide: .github/skills/eval-authoring/README.md — naming, placement, requirements, glossary, graders, anti-patterns, worked examples, local commands. Shared with eval-authoring-skill/eval-authoring-tool — update it there, not per-skill.
  • Repository-local eval README/configuration discovered in Step 0 (if present)

© Azure, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in .github/skills/eval-authoring-workflow of Azure/azure-sdk-tools.

  • SKILL.md
  • evals/eval.yaml

Open the folder on GitHubat commit 6e4fb2b

Compare with similar skills

Eval Authoring Workflow next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Eval Authoring Workflow compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Eval Authoring Workflow this skillAzure/azure-sdk-tools134—~972Automated safety check: PassMIT
Paperclip Evalspaperclipai/paperclip99k—~839Automated safety check: PassMIT
Run Evalslycorp-jp/sim-use1.4k—~1.4kAutomated safety check: PassApache-2.0
Run EvalsDevin-AXIS/iPolloWork6.8k—~851Automated safety check: PassCustom licence
Write A Specdifferent-ai/openwork24k—~3.3kAutomated safety check: PassCustom licence
Phoenix Pxi PlaywrightArize-ai/phoenix12k—~2.6kAutomated safety check: PassCustom licence

Similar skills

  • Paperclip Evals

    paperclipai/paperclip

    Choose, inspect, validate, and report Paperclip Runner or Product E2E evaluations while preserving evidence, provenance, cost, and failure classification.

    99k GitHub stars~839 tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Run Evals

    lycorp-jp/sim-use

    Prepare the environment and run the LLM-driven agent evals (e2e/agent-evals/) against a chosen sim-use binary.

    1.4k GitHub stars~1.4k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Run Evals

    Devin-AXIS/iPolloWork

    do e2e tests, run e2e, validate feature, prove it works, PR proof, frame proof, pnpm evals.

    6.8k GitHub stars~851 tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Write A Spec

    different-ai/openwork

    Write or extend an E2E journey spec in evals/specs that proves a PR's change to a human reviewer.

    24k GitHub stars~3.3k tokensUpdated today
    Testing & QAAuto-check passed
  • Phoenix Pxi Playwright

    Arize-ai/phoenix

    Write, extend, and debug PXI Playwright E2E tests for Phoenix.

    12k GitHub stars~2.6k tokensUpdated today
    Testing & QAAuto-check passed
  • LLM Benchmarking with lm-evaluation-harness

    Orchestra-Research/AI-Research-SKILLs

    Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.

    13k GitHub starsUsed in 8 repos~3k tokens
    AI & LLM EngineeringAuto-check passed

More from Azure/azure-sdk-tools

All 35 skills in this repo
  • Apiview Feedback Resolution

    Azure/azure-sdk-tools

    Official

    Analyze and resolve APIView review feedback on Azure SDK PRs.

    134 GitHub stars~547 tokensUpdated today
    Auto-check passed
  • Official

    Deploy test resources and run Azure SDK tests in live, record, or playback mode.

    134 GitHub stars~1.5k tokensUpdated today
    Auto-check: notes
  • Azsdk Common Pipeline Analysis

    Azure/azure-sdk-tools

    Official

    Analyze Azure SDK CI/CD pipeline failures into a structured diagnosis, and define the required output format.

    134 GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Official

    Create, get, update, abandon, and link SDK PRs to release plan work items for Azure SDK releases.

    134 GitHub stars~2.7k tokensUpdated today
    Auto-check passed
  • Azure Typespec Assessment

    Azure/azure-sdk-tools

    Official

    Assess Azure TypeSpec Git diffs for semantic intent, REST and downstream SDK breaking changes, Azure Guidelines compliance, and documentation completeness.

    134 GitHub stars~1.9k tokensUpdated today
    Auto-check passed

Questions about Eval Authoring Workflow

What does Eval Authoring Workflow do?

Author and validate multi-tool, multi-turn, mock, and live Vally scenarios under evals/workflows. Eval Authoring Workflow is an agent skill from Azure/azure-sdk-tools, published by the product's own GitHub organization. Author and validate multi-tool, multi-turn, mock, and live Vally scenarios under evals/workflows.

When should I use Eval Authoring Workflow?

Eval Authoring Workflow fits situations like: : per-skill routing/capability evals (use eval-authoring-skill); isolated prompt-to-tool tests (use eval-authoring-tool).

How do I install Eval Authoring Workflow in Claude Code?

Run `npx skills add Azure/azure-sdk-tools --skill eval-authoring-workflow -a claude-code`. Or copy the skill folder (.github/skills/eval-authoring-workflow in Azure/azure-sdk-tools) into .claude/skills/eval-authoring-workflow in your project. Claude Code loads it when a task matches its description.

How do I install Eval Authoring Workflow in Codex?

Run `npx skills add Azure/azure-sdk-tools --skill eval-authoring-workflow -a codex`. Or copy the skill folder (.github/skills/eval-authoring-workflow in Azure/azure-sdk-tools) into .agents/skills/eval-authoring-workflow in your project. Codex loads it when a task matches its description.

Can I use Eval Authoring Workflow in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Azure/azure-sdk-tools --skill eval-authoring-workflow -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/eval-authoring-workflow, .gemini/skills/eval-authoring-workflow, .github/skills/eval-authoring-workflow and .opencode/skills/eval-authoring-workflow in your project.

What does Eval Authoring Workflow need to run?

SKILL.md names no scripts, command-line tools or credentials: Eval Authoring Workflow is instructions for the agent only. Compatibility (from SKILL.md): copilot-chat, @microsoft/vally-cli 0.14.0.

Does Eval Authoring Workflow access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Eval Authoring Workflow safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Eval Authoring Workflow use?

Eval Authoring Workflow is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Eval Authoring Workflow use?

About 972 tokens (SKILL.md is roughly 3.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Eval Authoring Workflow?

Skills that share tags, products or a category with Eval Authoring Workflow: Paperclip Evals (paperclipai/paperclip, 99k stars), Run Evals (lycorp-jp/sim-use, 1.4k stars), Run Evals (Devin-AXIS/iPolloWork, 6.8k stars) and Write A Spec (different-ai/openwork, 24k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Eval Authoring Workflow?

Azure (a GitHub organization, an official publisher) maintains it in Azure/azure-sdk-tools, which has 134 GitHub stars. The repository holds 35 skills in this directory. The repository was last updated on October 8, 2026.

Source: Azure/azure-sdk-tools on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.