Official agent skill

SDK AI Bot Eval Dataset

by Azure in Azure/azure-sdk-tools

Create a new evaluation dataset or add cases to an existing one for the Azure SDK QA bot evaluation.

OfficialMITAuto-check: notesKnowledge Management

Install SDK AI Bot Eval Dataset

skills CLI
$ npx skills add Azure/azure-sdk-tools --skill sdk-ai-bot-eval-dataset -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Azure/azure-sdk-tools sdk-ai-bot-eval-dataset --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Azure/azure-sdk-tools.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.github/skills/sdk-ai-bot-eval-dataset .claude/skills/sdk-ai-bot-eval-dataset && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
sdk-ai-bot-eval-dataset
GitHub stars
134
Token cost
~1.4k tokens
SKILL.md length
439 words
Files
2 (incl. references)
Skills in repo
35
Repo updated
First seen
Licence
MIT

At a glance

Create a new evaluation dataset or add cases to an existing one for the Azure SDK QA bot evaluation.

  • Works in 5 steps: Pick a workflow from the table above… → Add or stage canonical rows in… → For staged cases, promote the reviewed… → …
  • : running evaluations
  • SKILL.md covers Triggers, Rules, Environment and Choose a workflow, plus 2 more sections
  • Calls python and az

What it does

SDK AI Bot Eval Dataset is an agent skill from Azure/azure-sdk-tools, published by the product's own GitHub organization. Create a new evaluation dataset or add cases to an existing one for the Azure SDK QA bot evaluation. WHEN: "add eval dataset item", "add a test case", "new evaluation dataset", "create dataset", "add question to dataset", "curate eval data", "promote staging cases", "upload dataset asset", "new scenario dataset". DO NOT USE FOR: running evaluations, pipeline troubleshooting, knowledge-graph indexing.

Its SKILL.md is about 1.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/dataset-schema-and-workflows.md`). Compatibility notes: local azure-sdk-tools clone, python 3.12 venv, az login

It sits in Knowledge Management, covering Test generation and Knowledge graphs. It works with Microsoft Azure. The repository describes itself as: Tools repository leveraged by the Azure SDK team. The licence is MIT.

When your agent uses it

  • : running evaluations
  • Pipeline troubleshooting
  • Knowledge-graph indexing

Example prompts

  • “add eval dataset item”
  • “add a test case”
  • “new evaluation dataset”
  • “/sdk-ai-bot-eval-dataset”

Requirements

  • Python 3
  • Compatibility (from SKILL.md): local azure-sdk-tools clone, python 3.12 venv, az login

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Pick a workflow from the table above (manual add, curate from blob, or new dataset).
  2. Add or stage canonical rows in evaluation_datasets//.jsonl.
  3. For staged cases, promote the reviewed ones with python -m dataset.review.
  4. Validate the file with python -m dataset.validate ... --require-reviewed.
  5. Publish with python -m dataset.upload, then commit the per-scenario file,

What it can do on your machine

Read from SKILL.md and the folder at commit cc5ca24. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python
    • az

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use az, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    local azure-sdk-tools clone, python 3.12 venv, az login

    From compatibility in the SKILL.md frontmatter.

Context cost

SDK AI Bot Eval Dataset loads about 1.4k tokens when it runs, and up to ~2.5k if it reads all its reference files. Until then it costs about 107 tokens; SKILL.md has 439 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~107
When it runs · the whole SKILL.md, loaded when a task matches
~1.4k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~2.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:38
    figure them. Dataset prep loads a local `.env` (copy and fill
  • NoteMentions a .env fileSKILL.md:49
    `.env` or the shell and re-run. A purely **manual add** (edit JSONL + validate) needs no

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Azure/azure-sdk-tools at commit cc5ca24, republished under its MIT licence (© Azure). 439 words, ~1,393 tokens.

Download SKILL.mdSave it as .claude/skills/sdk-ai-bot-eval-dataset/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
sdk-ai-bot-eval-dataset
description
Create a new evaluation dataset or add cases to an existing one for the Azure SDK QA bot evaluation. WHEN: "add eval dataset item", "add a test case", "new evaluation dataset", "create dataset", "add question to dataset", "curate eval data", "promote staging cases", "upload dataset asset", "new scenario dataset". DO NOT USE FOR: running evaluations, pipeline troubleshooting, knowledge-graph indexing.
compatibility
local azure-sdk-tools clone, python 3.12 venv, az login
license
MIT
metadata.version
1.0.0
metadata.distribution
local

QA Bot Evaluation Dataset

Create a new per-scenario evaluation dataset or add cases to an existing one for the QA bot evaluation package at tools/sdk-ai-bots/azure-sdk-qa-bot-evaluation. Datasets are per-scenario JSONL files under evaluation_datasets/<target>/<scenario>.jsonl (target = basic or perf) holding inputs + expectations only.

Run all commands from tools/sdk-ai-bots/azure-sdk-qa-bot-evaluation with the package .venv active and az login done. See schema and workflows for the canonical row format and step-by-step recipes.

Triggers

USE FOR: create a new evaluation dataset (new scenario file); add cases to an existing per-scenario dataset; curate cases from storage markdown; promote reviewed staging cases; upload a dataset as a Foundry asset WHEN: "add eval dataset item", "add a test case", "new evaluation dataset", "create dataset", "add question to dataset", "curate eval data", "promote staging cases", "upload dataset asset", "new scenario dataset" DO NOT USE FOR: running evaluations, pipeline troubleshooting, knowledge-graph indexing

Rules

  • A dataset is one file: evaluation_datasets/<target>/<scenario>.jsonl. Creating a new dataset = creating a new <scenario>.jsonl in basic/ or perf/.
  • The canonical dedup key is the normalized query (applied at curation). testcase titles may legitimately repeat (e.g. Untitled) — never dedup or fail on testcase.
  • Only reviewed: "pass" rows are curated/committed; see the review status lifecycle for the three states and how leftovers are finalized.
  • evaluation_datasets/_staging/ is committed (shared review state) so concurrent contributors don't re-curate the same cases; basic/, perf/ and registry.json are committed too.
  • Always validate before upload, and after editing any curated file.
Show full SKILL.md (205 more words)Show less

Environment

Before running any command that touches Azure, ensure the required variables are set and remind the user to configure them. Dataset prep loads a local .env (copy and fill in tools/sdk-ai-bots/azure-sdk-qa-bot-evaluation/env-variables) and authenticates with az login.

CommandRequires
dataset.curateaz login, STORAGE_BLOB_ACCOUNT, AI_ONLINE_PERFORMANCE_EVALUATION_STORAGE_CONTAINER
dataset.uploadaz login, AZURE_AI_PROJECT_ENDPOINT
dataset.validate, dataset.reviewnone (local file operations)

If a required variable is missing the command fails (KeyError / auth error) — set it in .env or the shell and re-run. A purely manual add (edit JSONL + validate) needs no env vars; only dataset.upload then requires AZURE_AI_PROJECT_ENDPOINT + az login.

Choose a workflow

GoalWorkflow
Add a few specific cases you already haveManual add
Harvest new cases from collected storage markdownCurate from blob
Create a brand-new scenario datasetNew dataset

Core commands

bash
# Validate a file or folder (--require-reviewed gates official datasets on reviewed=="pass")
python -m dataset.validate evaluation_datasets/<target>/<scenario>.jsonl --require-reviewed

# Promote reviewed (pass) staging rows; leftover items are finalized to abandoned
python -m dataset.review --target <basic|perf> [--scenario <scenario>]

# Upload one versioned Foundry asset per scenario; writes registry.json
python -m dataset.upload --target <basic|perf> [--scenario <scenario>]

After adding or creating a dataset: validate → upload → commit the per-scenario file, registry.json, and updated _staging/ files.

Steps

  1. Pick a workflow from the table above (manual add, curate from blob, or new dataset).
  2. Add or stage canonical rows in evaluation_datasets/<target>/<scenario>.jsonl.
  3. For staged cases, promote the reviewed ones with python -m dataset.review.
  4. Validate the file with python -m dataset.validate ... --require-reviewed.
  5. Publish with python -m dataset.upload, then commit the per-scenario file, registry.json, and updated _staging/.

© Azure, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (references) in .github/skills/sdk-ai-bot-eval-dataset of Azure/azure-sdk-tools.

  • SKILL.md
  • references/dataset-schema-and-workflows.md

Open the folder on GitHubat commit cc5ca24

Compare with similar skills

SDK AI Bot Eval Dataset next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

SDK AI Bot Eval Dataset compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
SDK AI Bot Eval Dataset this skillAzure/azure-sdk-tools134—~1.4kAutomated safety check: NotesMIT
Obsidian Canvas BoardsAgriciDaniel/claude-obsidian15k—~1.4kAutomated safety check: PassMIT
Ontology1mancompany/OneManCompany4412 repos~1.5kAutomated safety check: PassApache-2.0
Knowledge Graphgnomeria/usbtree691—~1.5kAutomated safety check: PassMIT
Graphagenticnotetaking/arscontexta3.5k—~4.9kAutomated safety check: NotesMIT
LLM Wiki Knowledge GraphEgonex-AI/Understand-Anything86k—~1.5kAutomated safety check: PassMIT

Similar skills

  • Obsidian Canvas Boards

    AgriciDaniel/claude-obsidian

    Creates, inspects and updates Obsidian JSON Canvas boards in a vault, with text, file, link, group and edge nodes, using safe recoverable edits.

    15k GitHub stars~1.4k tokensUpdated 29 days ago
    Knowledge ManagementAuto-check passed
  • Ontology

    1mancompany/OneManCompany

    Typed knowledge graph for structured agent memory and composable skills.

    441 GitHub starsUsed in 2 repos~1.5k tokens
    Knowledge ManagementAuto-check passed
  • Knowledge Graph

    gnomeria/usbtree

    Set up and maintain a lightweight, file-based knowledge graph of the repo — entities, typed relations, decisions, gotchas — so agents load context fast instead of re-exploring the codebase every…

    691 GitHub stars~1.5k tokensUpdated 1 mo ago
    Knowledge ManagementAuto-check passed
  • Graph

    agenticnotetaking/arscontexta

    Interactive knowledge graph analysis. An agent skill from agenticnotetaking/arscontexta.

    3.5k GitHub stars~4.9k tokensUpdated 7 mo ago
    Knowledge ManagementAuto-check: notes
  • LLM Wiki Knowledge Graph

    Egonex-AI/Understand-Anything

    Detects a Karpathy-pattern LLM wiki and builds an interactive knowledge graph with entities, implicit relationships and topic clusters.

    86k GitHub stars~1.5k tokensUpdated yesterday
    Knowledge ManagementAuto-check passed
  • Gitnexus Guide

    aws-samples/sample-kolya-br-proxy

    Official

    A skill your agent uses when the user asks about GitNexus itself — available tools, how to query the knowledge graph, MCP resources, graph schema, or workflow reference.

    106 GitHub starsUsed in 11 repos~867 tokens
    Knowledge ManagementAuto-check passed

More from Azure/azure-sdk-tools

All 35 skills in this repo
  • Apiview Feedback Resolution

    Azure/azure-sdk-tools

    Official

    Analyze and resolve APIView review feedback on Azure SDK PRs.

    134 GitHub stars~547 tokensUpdated today
    Auto-check passed
  • Official

    Deploy test resources and run Azure SDK tests in live, record, or playback mode.

    134 GitHub stars~1.5k tokensUpdated today
    Auto-check: notes
  • Azsdk Common Pipeline Analysis

    Azure/azure-sdk-tools

    Official

    Analyze Azure SDK CI/CD pipeline failures into a structured diagnosis, and define the required output format.

    134 GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Official

    Create, get, update, abandon, and link SDK PRs to release plan work items for Azure SDK releases.

    134 GitHub stars~2.7k tokensUpdated today
    Auto-check passed
  • Azure Typespec Assessment

    Azure/azure-sdk-tools

    Official

    Assess Azure TypeSpec Git diffs for semantic intent, REST and downstream SDK breaking changes, Azure Guidelines compliance, and documentation completeness.

    134 GitHub stars~1.9k tokensUpdated today
    Auto-check passed

Works with

Questions about SDK AI Bot Eval Dataset

What does SDK AI Bot Eval Dataset do?

Create a new evaluation dataset or add cases to an existing one for the Azure SDK QA bot evaluation. SDK AI Bot Eval Dataset is an agent skill from Azure/azure-sdk-tools, published by the product's own GitHub organization. Create a new evaluation dataset or add cases to an existing one for the Azure SDK QA bot evaluation.

When should I use SDK AI Bot Eval Dataset?

SDK AI Bot Eval Dataset fits situations like: : running evaluations; pipeline troubleshooting; knowledge-graph indexing.

How do I install SDK AI Bot Eval Dataset in Claude Code?

Run `npx skills add Azure/azure-sdk-tools --skill sdk-ai-bot-eval-dataset -a claude-code`. Or copy the skill folder (.github/skills/sdk-ai-bot-eval-dataset in Azure/azure-sdk-tools) into .claude/skills/sdk-ai-bot-eval-dataset in your project. Claude Code loads it when a task matches its description.

How do I install SDK AI Bot Eval Dataset in Codex?

Run `npx skills add Azure/azure-sdk-tools --skill sdk-ai-bot-eval-dataset -a codex`. Or copy the skill folder (.github/skills/sdk-ai-bot-eval-dataset in Azure/azure-sdk-tools) into .agents/skills/sdk-ai-bot-eval-dataset in your project. Codex loads it when a task matches its description.

Can I use SDK AI Bot Eval Dataset in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Azure/azure-sdk-tools --skill sdk-ai-bot-eval-dataset -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/sdk-ai-bot-eval-dataset, .gemini/skills/sdk-ai-bot-eval-dataset, .github/skills/sdk-ai-bot-eval-dataset and .opencode/skills/sdk-ai-bot-eval-dataset in your project.

What does SDK AI Bot Eval Dataset need to run?

Going by SKILL.md and its folder, SDK AI Bot Eval Dataset needs the command-line tools its instructions call (python and az). Our summary lists: Python 3. Compatibility (from SKILL.md): local azure-sdk-tools clone, python 3.12 venv, az login.

Does SDK AI Bot Eval Dataset access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is SDK AI Bot Eval Dataset safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does SDK AI Bot Eval Dataset use?

SDK AI Bot Eval Dataset is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does SDK AI Bot Eval Dataset use?

About 1.4k tokens (SKILL.md is roughly 5.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.1k tokens, read only when the agent opens those files.

What are the alternatives to SDK AI Bot Eval Dataset?

Skills that share tags, products or a category with SDK AI Bot Eval Dataset: Obsidian Canvas Boards (AgriciDaniel/claude-obsidian, 15k stars), Ontology (1mancompany/OneManCompany, 441 stars), Knowledge Graph (gnomeria/usbtree, 691 stars) and Graph (agenticnotetaking/arscontexta, 3.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains SDK AI Bot Eval Dataset?

Azure (a GitHub organization, an official publisher) maintains it in Azure/azure-sdk-tools, which has 134 GitHub stars. The repository holds 35 skills in this directory. The repository was last updated on October 9, 2026.

Source: Azure/azure-sdk-tools on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.