Azure AI Projects Python SDK
microsoft/skills
Reference for building on Microsoft Foundry with the azure-ai-projects Python SDK: project clients, versioned agents, evaluations, connections, datasets and indexes.
Convert an evaluation dataset or leaderboard into the Every Eval Ever (EEE) schema — aggregate .json logs (eval.schema.json) and optional instance samples.jsonl sidecars…
$ npx skills add evaleval/every_eval_ever --skill eee-dataset-conversion -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install evaleval/every_eval_ever eee-dataset-conversion --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/evaleval/every_eval_ever.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/eee-dataset-conversion .claude/skills/eee-dataset-conversion && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "eee-dataset-conversion" agent skill from https://github.com/evaleval/every_eval_ever/tree/main/.agents/skills/eee-dataset-conversion into .claude/skills/eee-dataset-conversion/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eee-dataset-conversion", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/evaleval/every_eval_ever/tree/main/.agents/skills/eee-dataset-conversionType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add evaleval/every_eval_ever --skill eee-dataset-conversion -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install evaleval/every_eval_ever eee-dataset-conversion --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/evaleval/every_eval_ever.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.agents/skills/eee-dataset-conversion .agents/skills/eee-dataset-conversion && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "eee-dataset-conversion" agent skill from https://github.com/evaleval/every_eval_ever/tree/main/.agents/skills/eee-dataset-conversion into .agents/skills/eee-dataset-conversion/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eee-dataset-conversion", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add evaleval/every_eval_ever --skill eee-dataset-conversion -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install evaleval/every_eval_ever eee-dataset-conversion --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/evaleval/every_eval_ever.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.agents/skills/eee-dataset-conversion .cursor/skills/eee-dataset-conversion && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "eee-dataset-conversion" agent skill from https://github.com/evaleval/every_eval_ever/tree/main/.agents/skills/eee-dataset-conversion into .cursor/skills/eee-dataset-conversion/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eee-dataset-conversion", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/evaleval/every_eval_ever.git --path .agents/skills/eee-dataset-conversion--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add evaleval/every_eval_ever --skill eee-dataset-conversion -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install evaleval/every_eval_ever eee-dataset-conversion --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/evaleval/every_eval_ever.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.agents/skills/eee-dataset-conversion .gemini/skills/eee-dataset-conversion && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "eee-dataset-conversion" agent skill from https://github.com/evaleval/every_eval_ever/tree/main/.agents/skills/eee-dataset-conversion into .gemini/skills/eee-dataset-conversion/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eee-dataset-conversion", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install evaleval/every_eval_ever eee-dataset-conversionInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add evaleval/every_eval_ever --skill eee-dataset-conversion -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/evaleval/every_eval_ever.git skills-src && mkdir -p .github/skills && cp -r skills-src/.agents/skills/eee-dataset-conversion .github/skills/eee-dataset-conversion && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "eee-dataset-conversion" agent skill from https://github.com/evaleval/every_eval_ever/tree/main/.agents/skills/eee-dataset-conversion into .github/skills/eee-dataset-conversion/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eee-dataset-conversion", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add evaleval/every_eval_ever --skill eee-dataset-conversion -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install evaleval/every_eval_ever eee-dataset-conversion --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/evaleval/every_eval_ever.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.agents/skills/eee-dataset-conversion .opencode/skills/eee-dataset-conversion && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "eee-dataset-conversion" agent skill from https://github.com/evaleval/every_eval_ever/tree/main/.agents/skills/eee-dataset-conversion into .opencode/skills/eee-dataset-conversion/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eee-dataset-conversion", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
eee-dataset-conversionConvert an evaluation dataset or leaderboard into the Every Eval Ever (EEE) schema — aggregate .json logs (eval.schema.json) and optional instance samples.jsonl sidecars…
Eee Dataset Conversion is an agent skill from evaleval/every_eval_ever. Convert an evaluation dataset or leaderboard into the Every Eval Ever (EEE) schema — aggregate .json logs (eval.schema.json) and optional instance samples.jsonl sidecars (instanceleveleval.schema.json). Use when asked to write an EEE adapter, add a dataset/leaderboard to the EEE datastore, map benchmark results into EEE, or debug why an EEE record won't validate.
Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 13 other files, including scripts (for example `reference/datastore-gate.md`, `reference/datastore-submission.md` and `reference/fields.md`).
It sits in AI & LLM Engineering, covering LLM evaluation. The repository describes itself as: Every Eval Ever is a shared schema and crowdsourced eval database. It defines a standardized metadata format for storing AI evaluation results — from leaderboard scrapes and… The licence is MIT.
7 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 1eb9d39. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/ (Python and Shell), which the agent can run.
Shell commands in SKILL.md call:
uvFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Eee Dataset Conversion loads about 2.5k tokens when it runs. Until then it costs about 99 tokens; SKILL.md has 1,224 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from evaleval/every_eval_ever at commit 1eb9d39, republished under its MIT licence (© evaleval). 1,224 words, ~2,480 tokens.
.claude/skills/eee-dataset-conversion/SKILL.md (or your agent's skills folder). This skill also uses 10 other files; get the full folder from GitHub.Rule you will keep relearning: all records must validate, but validating ≠ correct. Most real defects (answer leakage, double-counted aggregates, hardcoded scorers, non-idempotent ids) pass the schema and are still wrong. Always spot-check content, not just validity.
Written against EEE SCHEMA_VERSION 0.3.0 (import it from
every_eval_ever.helpers; never hardcode). If that value has moved, re-verify the
field claims in reference/ against the live schema — the schema always wins.
tests/test_skill_conversion.py pins this marker and re-validates this skill's
templates + frozen reference records, so a schema or validator change fails CI here
rather than in your PR. If that test is red, fix the skill, then regenerate the
frozen records.
How this runs. A person (the operator) runs you and can answer questions mid-run — you are not fully autonomous. When a choice sets policy (step 7's ask-list), ask the operator instead of deciding silently. Decide and log everything else. Finish with a PR that is ready to merge yet makes every non-obvious decision visible, so the maintainer who reviews it can comment and the skill/schema can improve. Two humans: the operator gates live; the PR informs the maintainer.
A source has model×benchmark scores (a leaderboard, a paper table, an HF results dataset, a harness dump) and you must emit EEE records. Two artifacts:
.json — one EvaluationLog per model (or per model×benchmark),
holding the headline scores. Always produced._samples.jsonl — one record per example. Only if you have
per-item data and want it.An adapter reshapes information the source already carries into the schema's form; it never manufactures information the source lacks.
1.0/0.0,
resolving a model string to its registry canonical_id, normalizing a metric
name, deriving the output dir from the collection. Reshaping a value you hold is
always fine, and the aggregate and instance-level paths must reshape the same
value the same way, so a result and its per-sample rows agree.bounds_status: unknown, a dropped-row failure), never
fill it in. This is the "fill" half of the coverage-vs-fill rule in
reference/fields.md §sources.reference/fields.md §sources
for the surface checklist, the coverage-vs-fill split, and which wins when they
disagree. Filter hygiene junk (.ipynb_checkpoints, *-checkpoint.json) and
segregate hand-curated baselines from harness runs.source_type is set by the artifact you hold, not who
ran the compute: raw per-item outputs → evaluation_run (even if a third party ran
them); only-aggregate reported numbers → documentation (a leaderboard scrape stays
documentation). Then: aggregate-only vs +instances; grain (one log per model =
default, or per model×benchmark when a benchmark has its own instance sidecar). See
reference/fields.md §shape.templates/aggregate_adapter.py
(always) and, for per-item data, templates/instance_sidecar.py (runnable
skeletons verified against the live validator). For a fuller real example,
mirror every_eval_ever/adapters/llm_stats (aggregate/documentation),
.../hfopenllm_v2 (documentation, many models), or .../openeval (aggregate +
instance sidecars). Adapters live at
every_eval_ever/adapters/<name>/adapter.py and run as
uv run python -m every_eval_ever.adapters.<name>.adapter; __init__.py just
marks the package. Don't hand-roll the write path or the drop path — the repo
owns both: publish through save_evaluation_logs (aggregate-only) or
converters.common.publication.publish_evaluation_logs (with instance sidecars), and
account for every rejected row via SourceConversionResult + save_failure_report +
a non-zero exit. See reference/datastore-gate.md §publish.reference/fields.md (aggregate) and reference/instance-level.md (jsonl).canonical_id for the join-key fields, with an opt-out
flag + never-fatal fallback to the raw id (marked unverified). But never key
evaluation_id on the resolved id — that's a moving join key; the record identity
rides the raw source id. See reference/registry.md.uv run python -m every_eval_ever validate <files> (files/glob, not a
dir), an offline unit test, ruff, a live smoke run, and a content spot-check.
The validator's semantic checks run only on the CLI, and only when the file sits at
its final data/<collection>/<dev>/<model>/ path. They are the merge gate, listed in
reference/datastore-gate.md. See reference/verification.md.skill-labeled PR or a
skill-gap issue — you needn't know the fix; flagging where you guessed is enough.source_data for a DB dump).yes (→ skill/skill-gap PR/issue) or no (dataset-specific).| Read this | When |
|---|---|
reference/fields.md | Filling any aggregate field; "which of the 3 source_* / 3 *_name fields?" |
reference/instance-level.md | Emitting _samples.jsonl: required fields, the interaction_type XOR, sample_hash, answer_attribution, the sidecar write-order |
reference/gotchas.md | Something validates but looks wrong; inf, double-counting, CI optional-deps, big-parquet reads |
reference/registry.md | Model/benchmark ids won't resolve; adding aliases |
reference/datastore-gate.md | What the CLI/bot enforce beyond the schema: paths, UUID4 names, companion pairing, score bounds, deployment axes, publishing |
reference/datastore-submission.md | Opening/updating the HF datastore PR: batching, the /eee validate bot, iterating without opening a new PR |
reference/verification.md | Before opening a PR; the checklist |
every_eval_ever/adapters/<name>/adapter.py +
__init__.py, a README.md (recommended), tests/test_<name>_adapter.py, + a row
in every_eval_ever/adapters/README.md). Code only — no generated records here.eval-card-registry repo (aliases / new canonicals) —
see its own CONTRIBUTING.md and the registry-entity-aliases skill there.EEE_datastore HF dataset (data/<collection>/, via a PR
with HfApi().upload_folder(..., create_pr=True)) — see
reference/datastore-submission.md for batching and the review bot.
Cross-link them. Reviewers ask for the adapter whenever data arrives without it, so
open the code PR even when the conversion was a one-off script.Schemas are the source of truth — when a reference and the schema disagree, the
schema wins; read eval.schema.json / instance_level_eval.schema.json.
© evaleval, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 10 other files (scripts) in .agents/skills/eee-dataset-conversion of evaleval/every_eval_ever.
Open the folder on GitHubat commit 1eb9d39
Eee Dataset Conversion next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Eee Dataset Conversion this skillevaleval/every_eval_ever | 133 | — | ~2.5k | Automated safety check: Pass | MIT | |
| Azure AI Projects Python SDKmicrosoft/skills | 3.1k | 6 repos | ~2.8k | Automated safety check: Pass | MIT | |
| Spec Optimizeleo-kuang-ai/spec-first | 107 | — | ~13k | Automated safety check: Pass | MIT | |
| Factory Learntikalk/adlc-team-skills | 141 | — | ~1.5k | Automated safety check: Pass | MIT | |
| Evals Contextzgsm-ai/costrict | 4.4k | 1 repos | ~1.9k | Automated safety check: Pass | Apache-2.0 | |
| Fine-Tuning ExpertJeffallan/claude-skills | 12k | 1 repos | ~1.7k | Automated safety check: Pass | MIT |
microsoft/skills
Reference for building on Microsoft Foundry with the azure-ai-projects Python SDK: project clients, versioned agents, evaluations, connections, datasets and indexes.
leo-kuang-ai/spec-first
Run metric-driven iterative optimization loops. An agent skill from leo-kuang-ai/spec-first.
tikalk/adlc-team-skills
A skill your agent uses when coordinating continuous improvement loops (team-levelup + change + evals feedback + cleanup) targeting team-ai-directives — includes build-to-delete pruning and…
zgsm-ai/costrict
Provides context about the CoStrict evals system structure in this monorepo.
Jeffallan/claude-skills
Guides LLM fine-tuning with LoRA and QLoRA through Hugging Face PEFT, from dataset validation and training checks to adapter merging, quantization and deployment.
huggingface/skills
Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.
evaleval/every_eval_ever
Review and repair pull requests on the evaleval/EEEdatastore Hugging Face dataset.
Categories
Convert an evaluation dataset or leaderboard into the Every Eval Ever (EEE) schema — aggregate .json logs (eval.schema.json) and optional instance samples.jsonl sidecars…. Eee Dataset Conversion is an agent skill from evaleval/every_eval_ever.json).
Eee Dataset Conversion fits situations like: asked to write an EEE adapter; add a dataset/leaderboard to the EEE datastore; map benchmark results into EEE; debug why an EEE record wont validate.
Run `npx skills add evaleval/every_eval_ever --skill eee-dataset-conversion -a claude-code`. Or copy the skill folder (.agents/skills/eee-dataset-conversion in evaleval/every_eval_ever) into .claude/skills/eee-dataset-conversion in your project. Claude Code loads it when a task matches its description.
Run `npx skills add evaleval/every_eval_ever --skill eee-dataset-conversion -a codex`. Or copy the skill folder (.agents/skills/eee-dataset-conversion in evaleval/every_eval_ever) into .agents/skills/eee-dataset-conversion in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add evaleval/every_eval_ever --skill eee-dataset-conversion -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/eee-dataset-conversion, .gemini/skills/eee-dataset-conversion, .github/skills/eee-dataset-conversion and .opencode/skills/eee-dataset-conversion in your project.
Going by SKILL.md and its folder, Eee Dataset Conversion needs Python and a shell for the scripts in its folder and the command-line tools its instructions call (uv). Our summary lists: Python 3; A Bash shell.
SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Eee Dataset Conversion is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.5k tokens (SKILL.md is roughly 9.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Eee Dataset Conversion: Azure AI Projects Python SDK (microsoft/skills, 3.1k stars), Spec Optimize (leo-kuang-ai/spec-first, 107 stars), Factory Learn (tikalk/adlc-team-skills, 141 stars) and Evals Context (zgsm-ai/costrict, 4.4k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
evaleval (a GitHub organization) maintains it in evaleval/every_eval_ever, which has 133 GitHub stars. The repository holds 2 skills in this directory. The repository was last updated on October 7, 2026.
Source: evaleval/every_eval_ever on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.