Official agent skill

Vally Eval

by microsoft in microsoft/GitHub-Copilot-for-Azure

Author, validate, and run Vally evaluation suites for agent skills.

OfficialMITAuto-check passedTesting & QA

Install Vally Eval

skills CLI
$ npx skills add microsoft/GitHub-Copilot-for-Azure --skill vally-eval -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install microsoft/GitHub-Copilot-for-Azure vally-eval --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/microsoft/GitHub-Copilot-for-Azure.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.github/skills/vally-eval .claude/skills/vally-eval && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
vally-eval
GitHub stars
255
Token cost
~1.5k tokens
SKILL.md length
743 words
Files
2 (incl. references)
Skills in repo
56
Repo updated
First seen
Licence
MIT

At a glance

Author, validate, and run Vally evaluation suites for agent skills.

  • Testing & QA work in your project
  • SKILL.md covers Write vally eval suites, Why is there a custom executor, Validate vally eval suites and Run vally eval suites locally, plus 3 more sections
  • Calls npm and npx

What it does

Vally Eval is an agent skill from microsoft/GitHub-Copilot-for-Azure, published by the product's own GitHub organization. Author, validate, and run Vally evaluation suites for agent skills. TRIGGERS: create eval, write eval, add eval, run eval, validate eval, vally eval, eval.yaml, add stimulus, map test to eval, migrate test to eval, eval graders, eval scoring, add eval to CI.

Its SKILL.md is about 1.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/ci-test.md`).

It sits in Testing & QA. It works with Microsoft Azure. The repository describes itself as: GitHub Copilot for Azure. The licence is MIT.

When your agent uses it

  • Testing & QA work in your project

Example prompts

  • “/vally-eval”

Requirements

  • Node.js

What it can do on your machine

Read from SKILL.md and the folder at commit d8f4f4e. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • npm
    • npx

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • microsoft.github.io
    • aka.ms

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Vally Eval loads about 1.5k tokens when it runs, and up to ~2.2k if it reads all its reference files. Until then it costs about 67 tokens; SKILL.md has 743 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~67
When it runs · the whole SKILL.md, loaded when a task matches
~1.5k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~2.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from microsoft/GitHub-Copilot-for-Azure at commit d8f4f4e, republished under its MIT licence (© microsoft). 743 words, ~1,544 tokens.

Download SKILL.mdSave it as .claude/skills/vally-eval/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
vally-eval
description
Author, validate, and run Vally evaluation suites for agent skills. TRIGGERS: create eval, write eval, add eval, run eval, validate eval, vally eval, eval.yaml, add stimulus, map test to eval, migrate test to eval, eval graders, eval scoring, add eval to CI.
license
MIT
metadata.author
Microsoft
metadata.version
1.0.0

Vally eval suites

Skills in the azure-skills plugin are required to have integration tests that run prompts against an LLM agent to evaluate whether they help the agent accomplish goals in target scenarios. Such integration tests are written as vally eval suites, using vally as the underlying tool for running tests and grading the agent outcome.

Write vally eval suites

Vally eval suites are written as yaml documents. All eval suites share eval spec.

Refer to the official documentation on the schema of the spec and the schema of the eval suites writing-eval-specs.

Vally eval suites for azure-skills plugin have the following file layout. The shared eval spec is located at <repo-root>/.vally.yaml. The eval suites are categorized by plugin and skills. The eval suites for each skill are located at <repo-root>/evals/<plugin-dirname>/<skill-name>/*.yaml, e.g. <repo-root>/evals/azure-skills/azure-ai/eval.yaml.

Use meaningful file names to categorize tests. If a skill needs fixture files for its eval suites, it should organize such fixture files in a fixture directory under its directory, e.g. <repo-root>/evals/azure-skills/azure-ai/fixture/. The vally test runner and stimulus validation script will load all *.yaml files except for those under a fixture/ directory. Make sure to put all fixture files under the fixture/ directory.

Why is there a custom executor

Our custom executor implemented features that vally doesn't support yet, such as early termination, system prompt modification, screenshot taking, etc. Besides, test-all-integration runs automated integration tests, collects its exported data and feeds the data to a dashboard web app under <repo-root>/dashboard/ to monitor skill integration test results.

If you intend to have your vally suites use any of the extended features or have their results be consumed by the dashboard, you MUST use the custom executor in your vally suites.

Use tags to control the custom executor

The custom executor uses special tag values to control the behavior of the custom executor. See tag-helpers.ts to learn what special tags are supported.

Note: If an eval suite specifies an earlyTerminate condition, the suite MUST NOT use the completed grader because early terminated runs will always fail the completed grader by design.

Validate vally eval suites

Vally eval suites in this repo follow certain conventions. For example, all eval suites must have a type, tier, cost and area tag so they can be run for a corresponding target group. To ensure all eval suites follow the conventions, a script is added to validate the eval suites and report errors when it sees any violation. To run the script, execute this command from the scripts/ directory.

bash
# cwd as <repo-root>/scripts/
npm run vally validate-stimulus

Extended features such as early termination are implemented using tags and many of them use serialized JSON objects as input. This validation script also validates the values of these special tags.

Show full SKILL.md (297 more words)Show less

Run vally eval suites locally

Use vally-cli to run vally eval suites. In most cases, you would like to use a command like this.

bash
# In tests/
npm run test:vally -- --plugin $PLUGIN_DIR --skill $SKILL

See vally test runner on how it composes the vally commands under the hood.

Run vally eval suites in CI

Vally eval suites implemented in this repo can be added to the CI test workflow to be run nightly and publish results for reviewing. Refer to ci-test on how to add the Vally eval suites to the CI test workflow.

Extend with custom grader

Custom graders can be added to grade trajectories in ways built-in graders don't support. To add a custom grader, follow the examples in the official vally documentation to create a tests/vally/<custom-name>-grader.ts module and register the new custom grader in tests/vally/vally-graders.ts. The npm run test:vally command internally loads all the custom graders when testing skills.

Re-grade an existing trajectory

Test authors commonly need to fine-tune grader configurations to reduce result flakiness. Vally supports re-grading an existing trajectory using a command like this:

bash
# in tests/
npx @microsoft/vally-cli grade --eval-spec ../evals/<plugin-dirname>/<skill-name>/eval.yaml --verbose < results/<test-run-name>/results.jsonl

You can keep tuning the grader config in eval.yaml and re-grade the trajectory until the results meet your expectations.

If your skill uses a custom grader, add --grader-plugin to load the custom graders.

bash
# in tests/
npx @microsoft/vally-cli grade --eval-spec ../evals/<plugin-dirname>/<skill-name>/eval.yaml --grader-plugin
../../../tests/vally/vally-graders.ts --verbose < results/<test-run-name>/results.jsonl

Note that the grader plugin's path is relative to the parent directory of the eval spec to run. For example, if the eval spec to run is <repo-root>/evals/azure-skills/azure-ai/eval.yaml, resolving this relative path ends at <repo-root>/tests/vally/vally-executor.ts.

Collect test results

When running locally, the test results can be found at the following directories:

  • tests/reports/<test-run-name>/
  • tests/results/<test-run-name>/

When running in CI, the test results can be found in the GitHub Action artifacts or at a storage account that the workflow publishes to. You can also use the integration tests dashboard to view the test results from nightly test runs.

© microsoft, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (references) in .github/skills/vally-eval of microsoft/GitHub-Copilot-for-Azure.

  • SKILL.md
  • references/ci-test.md

Open the folder on GitHubat commit d8f4f4e

Compare with similar skills

Vally Eval next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Vally Eval compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Vally Eval this skillmicrosoft/GitHub-Copilot-for-Azure255—~1.5kAutomated safety check: PassMIT
Run E2E Testkubernetes-sigs/cloud-provider-azure294—~3.8kAutomated safety check: PassApache-2.0
Dev ServerAzure/cosmos-explorer131—~1.6kAutomated safety check: PassMIT
WebGPU Provider Testing Without a GPUmicrosoft/onnxruntime22k—~1.4kAutomated safety check: PassMIT
Verify PRAzure/LogicAppsUX114—~2.4kAutomated safety check: PassMIT
Avm Tf TestingAzure/terraform-azurerm-avm-ptn-alz135—~1.8kAutomated safety check: PassMIT

Similar skills

  • Run E2E Test

    kubernetes-sigs/cloud-provider-azure

    Official

    Parse a Go e2e test from tests/e2e/, translate each step to kubectl and az CLI commands, and interactively replay the test against a live cluster.

    294 GitHub stars~3.8k tokensUpdated today
    Testing & QAAuto-check passed
  • Dev Server

    Azure/cosmos-explorer

    Official

    Start the local webpack dev server and connect to it with the Playwright browser.

    131 GitHub stars~1.6k tokensUpdated 2 days ago
    Testing & QAAuto-check passed
  • Official

    Shows how to build and run ONNX Runtime WebGPU provider tests on Linux with no GPU, using the Mesa lavapipe software Vulkan adapter, and where that approach falls short.

    22k GitHub stars~1.4k tokensUpdated today
    Testing & QAAuto-check passed
  • Verify PR

    Azure/LogicAppsUX

    Official

    Verify PR readiness before requesting review or merging. An agent skill from Azure/LogicAppsUX.

    114 GitHub stars~2.4k tokensUpdated today
    Testing & QAAuto-check passed
  • Avm Tf Testing

    Azure/terraform-azurerm-avm-ptn-alz

    Official

    A skill your agent uses for AVM Terraform validation, provider-mocked unit tests, real-Azure integration tests, E2E example tests, PowerShell hooks, OIDC, policy checks, and Avm.Authoring CI behavior.

    135 GitHub stars~1.8k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Azure Deployment

    EmeaAppGbb/spec2cloud

    Provision Azure infrastructure, deploy to Azure Container Apps, and verify via smoke tests.

    100 GitHub stars~1.8k tokensUpdated 5 mo ago
    Testing & QAAuto-check passed

More from microsoft/GitHub-Copilot-for-Azure

All 56 skills in this repo
  • Capacity

    microsoft/GitHub-Copilot-for-Azure

    Official

    Discovers available Azure OpenAI model capacity across regions and projects.

    255 GitHub starsUsed in 2 repos~1.7k tokens
    Auto-check passed
  • Deploy Model

    microsoft/GitHub-Copilot-for-Azure

    Official

    Unified Azure OpenAI model deployment skill with intelligent intent-based routing.

    255 GitHub starsUsed in 1 repo~1.8k tokens
    Auto-check passed
  • Entra Agent Id

    microsoft/GitHub-Copilot-for-Azure

    Official

    Provision Microsoft Entra Agent Identity Blueprints, BlueprintPrincipals, and per-instance Agent Identities via Microsoft Graph, and configure OAuth 2.0 token exchange (fmipath, OBO, cross-tenant)…

    255 GitHub starsUsed in 3 repos~4k tokens
    Auto-check passed
  • Microsoft Foundry

    microsoft/GitHub-Copilot-for-Azure

    Official

    Build, deploy, evaluate, optimize, fine-tune, and manage Microsoft Foundry agents, models, and resources end to end.

    255 GitHub starsUsed in 1 repo~6.7k tokens
    Auto-check passed
  • Azure Storage

    microsoft/GitHub-Copilot-for-Azure

    Official

    Azure Storage Services including Blob Storage, File Shares, Queue Storage, Table Storage, and Data Lake.

    255 GitHub starsUsed in 2 repos~1.3k tokens
    Auto-check passed
  • Azure Diagnostics

    microsoft/GitHub-Copilot-for-Azure

    Official

    Debug Azure production issues on Azure using AppLens, Azure Monitor, resource health, and safe triage.

    255 GitHub starsUsed in 1 repo~1.6k tokens
    Auto-check passed

Works with

Categories

Questions about Vally Eval

What does Vally Eval do?

Author, validate, and run Vally evaluation suites for agent skills. Vally Eval is an agent skill from microsoft/GitHub-Copilot-for-Azure, published by the product's own GitHub organization. Author, validate, and run Vally evaluation suites for agent skills.

When should I use Vally Eval?

Vally Eval fits situations like: testing & QA work in your project.

How do I install Vally Eval in Claude Code?

Run `npx skills add microsoft/GitHub-Copilot-for-Azure --skill vally-eval -a claude-code`. Or copy the skill folder (.github/skills/vally-eval in microsoft/GitHub-Copilot-for-Azure) into .claude/skills/vally-eval in your project. Claude Code loads it when a task matches its description.

How do I install Vally Eval in Codex?

Run `npx skills add microsoft/GitHub-Copilot-for-Azure --skill vally-eval -a codex`. Or copy the skill folder (.github/skills/vally-eval in microsoft/GitHub-Copilot-for-Azure) into .agents/skills/vally-eval in your project. Codex loads it when a task matches its description.

Can I use Vally Eval in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add microsoft/GitHub-Copilot-for-Azure --skill vally-eval -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/vally-eval, .gemini/skills/vally-eval, .github/skills/vally-eval and .opencode/skills/vally-eval in your project.

What does Vally Eval need to run?

Going by SKILL.md and its folder, Vally Eval needs the command-line tools its instructions call (npm and npx). Our summary lists: Node.js.

Does Vally Eval access the network?

SKILL.md names 2 domains. As links in the text: microsoft.github.io and aka.ms. This is read from the text; nothing was executed.

Is Vally Eval safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Vally Eval use?

Vally Eval is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Vally Eval use?

About 1.5k tokens (SKILL.md is roughly 6.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 658 tokens, read only when the agent opens those files.

What are the alternatives to Vally Eval?

Skills that share tags, products or a category with Vally Eval: Run E2E Test (kubernetes-sigs/cloud-provider-azure, 294 stars), Dev Server (Azure/cosmos-explorer, 131 stars), WebGPU Provider Testing Without a GPU (microsoft/onnxruntime, 22k stars) and Verify PR (Azure/LogicAppsUX, 114 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Vally Eval?

microsoft (a GitHub organization, an official publisher) maintains it in microsoft/GitHub-Copilot-for-Azure, which has 255 GitHub stars. The repository holds 56 skills in this directory. The repository was last updated on October 7, 2026.

Source: microsoft/GitHub-Copilot-for-Azure on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.