Agent skill

Eee Dataset Conversion

by evaleval in evaleval/every_eval_ever

Convert an evaluation dataset or leaderboard into the Every Eval Ever (EEE) schema — aggregate .json logs (eval.schema.json) and optional instance samples.jsonl sidecars…

MITAuto-check passedAI & LLM Engineering

Install Eee Dataset Conversion

skills CLI
$ npx skills add evaleval/every_eval_ever --skill eee-dataset-conversion -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install evaleval/every_eval_ever eee-dataset-conversion --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/evaleval/every_eval_ever.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/eee-dataset-conversion .claude/skills/eee-dataset-conversion && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
eee-dataset-conversion
GitHub stars
133
Token cost
~2.5k tokens
SKILL.md length
1,224 words
Files
11 (incl. scripts)
Skills in repo
2
Repo updated
First seen
Licence
MIT

At a glance

Convert an evaluation dataset or leaderboard into the Every Eval Ever (EEE) schema — aggregate .json logs (eval.schema.json) and optional instance samples.jsonl sidecars…

  • Works in 7 steps: Inspect the source first — you can't map… → Decide the shape — source_type is set by… → Copy a template / reference adapter —… → …
  • Asked to write an EEE adapter
  • SKILL.md covers When this skill applies, The one rule: canonicalize,…, Workflow (do these in order) and Load a reference only when you…, plus 1 more section
  • Runs Python and Shell scripts from its folder; calls uv

What it does

Eee Dataset Conversion is an agent skill from evaleval/every_eval_ever. Convert an evaluation dataset or leaderboard into the Every Eval Ever (EEE) schema — aggregate .json logs (eval.schema.json) and optional instance samples.jsonl sidecars (instanceleveleval.schema.json). Use when asked to write an EEE adapter, add a dataset/leaderboard to the EEE datastore, map benchmark results into EEE, or debug why an EEE record won't validate.

Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 13 other files, including scripts (for example `reference/datastore-gate.md`, `reference/datastore-submission.md` and `reference/fields.md`).

It sits in AI & LLM Engineering, covering LLM evaluation. The repository describes itself as: Every Eval Ever is a shared schema and crowdsourced eval database. It defines a standardized metadata format for storing AI evaluation results — from leaderboard scrapes and… The licence is MIT.

When your agent uses it

  • Asked to write an EEE adapter
  • Add a dataset/leaderboard to the EEE datastore
  • Map benchmark results into EEE
  • Debug why an EEE record wont validate

Example prompts

  • “/eee-dataset-conversion”

Requirements

  • Python 3
  • A Bash shell

Workflow steps

7 steps, taken from the first numbered list in SKILL.md.

  1. Inspect the source first — you can't map fields you haven't seen. Establish
  2. Decide the shape — source_type is set by the artifact you hold, not who
  3. Copy a template / reference adapter — templates/aggregate_adapter.py
  4. Fill fields carefully — the field traps are the whole game. Load
  5. Canonicalize ids — model + benchmark ids must resolve in the
  6. Verify — uv run python -m every_eval_ever validate (files/glob, not a
  7. Ask, then log your decisions. Two channels, don't confuse them

What it can do on your machine

Read from SKILL.md and the folder at commit 1eb9d39. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python and Shell), which the agent can run.

    Shell commands in SKILL.md call:

    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Eee Dataset Conversion loads about 2.5k tokens when it runs. Until then it costs about 99 tokens; SKILL.md has 1,224 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~99
When it runs · the whole SKILL.md, loaded when a task matches
~2.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from evaleval/every_eval_ever at commit 1eb9d39, republished under its MIT licence (© evaleval). 1,224 words, ~2,480 tokens.

Download SKILL.mdSave it as .claude/skills/eee-dataset-conversion/SKILL.md (or your agent's skills folder). This skill also uses 10 other files; get the full folder from GitHub.
name
eee-dataset-conversion
description
Convert an evaluation dataset or leaderboard into the Every Eval Ever (EEE) schema — aggregate `.json` logs (eval.schema.json) and optional instance `_samples.jsonl` sidecars (instance_level_eval.schema.json). Use when asked to write an EEE adapter, add a dataset/leaderboard to the EEE datastore, map benchmark results into EEE, or debug why an EEE record won't validate.
license
MIT
metadata.version
0.1.0

Converting evaluation results into Every Eval Ever (EEE)

Rule you will keep relearning: all records must validate, but validating ≠ correct. Most real defects (answer leakage, double-counted aggregates, hardcoded scorers, non-idempotent ids) pass the schema and are still wrong. Always spot-check content, not just validity.

Written against EEE SCHEMA_VERSION 0.3.0 (import it from every_eval_ever.helpers; never hardcode). If that value has moved, re-verify the field claims in reference/ against the live schema — the schema always wins. tests/test_skill_conversion.py pins this marker and re-validates this skill's templates + frozen reference records, so a schema or validator change fails CI here rather than in your PR. If that test is red, fix the skill, then regenerate the frozen records.

How this runs. A person (the operator) runs you and can answer questions mid-run — you are not fully autonomous. When a choice sets policy (step 7's ask-list), ask the operator instead of deciding silently. Decide and log everything else. Finish with a PR that is ready to merge yet makes every non-obvious decision visible, so the maintainer who reviews it can comment and the skill/schema can improve. Two humans: the operator gates live; the PR informs the maintainer.

When this skill applies

A source has model×benchmark scores (a leaderboard, a paper table, an HF results dataset, a harness dump) and you must emit EEE records. Two artifacts:

  • Aggregate .json — one EvaluationLog per model (or per model×benchmark), holding the headline scores. Always produced.
  • **Instance _samples.jsonl — one record per example. Only if you have per-item data and want it.

The one rule: canonicalize, never invent

An adapter reshapes information the source already carries into the schema's form; it never manufactures information the source lacks.

  • Canonicalization is expected — coercing a boolean pass/fail to 1.0/0.0, resolving a model string to its registry canonical_id, normalizing a metric name, deriving the output dir from the collection. Reshaping a value you hold is always fine, and the aggregate and instance-level paths must reshape the same value the same way, so a result and its per-sample rows agree.
  • Invention is not — a score for an unscored sample, bounds for a metric whose range is unknown, a model identity the source never named. Record the gap as absent (a null result id, bounds_status: unknown, a dropped-row failure), never fill it in. This is the "fill" half of the coverage-vs-fill rule in reference/fields.md §sources.

Workflow (do these in order)

  1. Inspect the source first — you can't map fields you haven't seen. Establish: distinct models · benchmarks/subtasks · the metric and its range · is there per-item data · the harness · timestamps · provenance (paper + each benchmark's own dataset repo). These facts are usually spread across many surfaces and which lives where varies per dataset, so gather every relevant surface before recording a field as unknown** — see reference/fields.md §sources for the surface checklist, the coverage-vs-fill split, and which wins when they disagree. Filter hygiene junk (.ipynb_checkpoints, *-checkpoint.json) and segregate hand-curated baselines from harness runs.
  2. Decide the shape — source_type is set by the artifact you hold, not who ran the compute: raw per-item outputs → evaluation_run (even if a third party ran them); only-aggregate reported numbers → documentation (a leaderboard scrape stays documentation). Then: aggregate-only vs +instances; grain (one log per model = default, or per model×benchmark when a benchmark has its own instance sidecar). See reference/fields.md §shape.
  3. Copy a template / reference adapter — templates/aggregate_adapter.py (always) and, for per-item data, templates/instance_sidecar.py (runnable skeletons verified against the live validator). For a fuller real example, mirror every_eval_ever/adapters/llm_stats (aggregate/documentation), .../hfopenllm_v2 (documentation, many models), or .../openeval (aggregate + instance sidecars). Adapters live at every_eval_ever/adapters/<name>/adapter.py and run as uv run python -m every_eval_ever.adapters.<name>.adapter; __init__.py just marks the package. Don't hand-roll the write path or the drop path — the repo owns both: publish through save_evaluation_logs (aggregate-only) or converters.common.publication.publish_evaluation_logs (with instance sidecars), and account for every rejected row via SourceConversionResult + save_failure_report + a non-zero exit. See reference/datastore-gate.md §publish.
  4. Fill fields carefully — the field traps are the whole game. Load reference/fields.md (aggregate) and reference/instance-level.md (jsonl).
  5. Canonicalize ids — model + benchmark ids must resolve in the eval-card-registry (else they fragment the data). Default: resolve live against the hosted resolver and use canonical_id for the join-key fields, with an opt-out flag + never-fatal fallback to the raw id (marked unverified). But never key evaluation_id on the resolved id — that's a moving join key; the record identity rides the raw source id. See reference/registry.md.
  6. Verify — uv run python -m every_eval_ever validate <files> (files/glob, not a dir), an offline unit test, ruff, a live smoke run, and a content spot-check. The validator's semantic checks run only on the CLI, and only when the file sits at its final data/<collection>/<dev>/<model>/ path. They are the merge gate, listed in reference/datastore-gate.md. See reference/verification.md.
  7. Ask, then log your decisions. Two channels, don't confuse them:
    • Ask the operator (live) when a choice sets policy: creating a new canonical id · dropping a non-trivial share of the data · an ambiguous metric choice · bounding an unbounded metric · re-hosting large data · the source won't fit without a structural change (a schema field, an edit to a base adapter or a shared converter, relaxing a validator rule). Don't decide these silently — the person running you is there to answer. The structural one is not yours to fold into this PR: its design gets agreed before a PR exists, and carrying it here would hold the adapter behind that discussion.
    • Log (in the PR) every non-obvious choice — not just where it was hard. A confident wrong choice produces no "friction," so log decisions, not pain. Finish with a ready-to-merge PR carrying the decision log below. General gaps (would recur on other datasets) also become a separate skill-labeled PR or a skill-gap issue — you needn't know the fix; flagging where you guessed is enough.
Show full SKILL.md (280 more words)Show less
Decision log (paste into the PR description)
  • Decision / where — the field or step (e.g. source_data for a DB dump).
  • Chose / instead of — what you did and the alternative you rejected.
  • Confidence — high / medium / low (low = please, maintainer, look here).
  • General? — yes (→ skill/skill-gap PR/issue) or no (dataset-specific).
  • Coverage (once per adapter) — "N source rows → M records, K dropped (reason)". No silent caps — if you filtered/sampled/capped anything, say so here.

Load a reference only when you need it (progressive disclosure)

Read thisWhen
reference/fields.mdFilling any aggregate field; "which of the 3 source_* / 3 *_name fields?"
reference/instance-level.mdEmitting _samples.jsonl: required fields, the interaction_type XOR, sample_hash, answer_attribution, the sidecar write-order
reference/gotchas.mdSomething validates but looks wrong; inf, double-counting, CI optional-deps, big-parquet reads
reference/registry.mdModel/benchmark ids won't resolve; adding aliases
reference/datastore-gate.mdWhat the CLI/bot enforce beyond the schema: paths, UUID4 names, companion pairing, score bounds, deployment axes, publishing
reference/datastore-submission.mdOpening/updating the HF datastore PR: batching, the /eee validate bot, iterating without opening a new PR
reference/verification.mdBefore opening a PR; the checklist

The three PRs a contribution usually is

  1. Adapter code → this repo (every_eval_ever/adapters/<name>/adapter.py + __init__.py, a README.md (recommended), tests/test_<name>_adapter.py, + a row in every_eval_ever/adapters/README.md). Code only — no generated records here.
  2. Canonical ids → the eval-card-registry repo (aliases / new canonicals) — see its own CONTRIBUTING.md and the registry-entity-aliases skill there.
  3. Generated data → the EEE_datastore HF dataset (data/<collection>/, via a PR with HfApi().upload_folder(..., create_pr=True)) — see reference/datastore-submission.md for batching and the review bot. Cross-link them. Reviewers ask for the adapter whenever data arrives without it, so open the code PR even when the conversion was a one-off script.

Schemas are the source of truth — when a reference and the schema disagree, the schema wins; read eval.schema.json / instance_level_eval.schema.json.

© evaleval, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 10 other files (scripts) in .agents/skills/eee-dataset-conversion of evaleval/every_eval_ever.

  • SKILL.md
  • reference/datastore-gate.md
  • reference/datastore-submission.md
  • reference/fields.md
  • reference/gotchas.md
  • reference/instance-level.md
  • reference/registry.md
  • reference/verification.md
  • scripts/validate.sh
  • templates/aggregate_adapter.py
  • templates/instance_sidecar.py

Open the folder on GitHubat commit 1eb9d39

Compare with similar skills

Eee Dataset Conversion next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Eee Dataset Conversion compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Eee Dataset Conversion this skillevaleval/every_eval_ever133—~2.5kAutomated safety check: PassMIT
Azure AI Projects Python SDKmicrosoft/skills3.1k6 repos~2.8kAutomated safety check: PassMIT
Spec Optimizeleo-kuang-ai/spec-first107—~13kAutomated safety check: PassMIT
Factory Learntikalk/adlc-team-skills141—~1.5kAutomated safety check: PassMIT
Evals Contextzgsm-ai/costrict4.4k1 repos~1.9kAutomated safety check: PassApache-2.0
Fine-Tuning ExpertJeffallan/claude-skills12k1 repos~1.7kAutomated safety check: PassMIT

Similar skills

  • Official

    Reference for building on Microsoft Foundry with the azure-ai-projects Python SDK: project clients, versioned agents, evaluations, connections, datasets and indexes.

    3.1k GitHub starsUsed in 6 repos~2.8k tokens
    AI & LLM EngineeringAuto-check passed
  • Spec Optimize

    leo-kuang-ai/spec-first

    Run metric-driven iterative optimization loops. An agent skill from leo-kuang-ai/spec-first.

    107 GitHub stars~13k tokensUpdated 13 days ago
    AI & LLM EngineeringAuto-check passed
  • Factory Learn

    tikalk/adlc-team-skills

    A skill your agent uses when coordinating continuous improvement loops (team-levelup + change + evals feedback + cleanup) targeting team-ai-directives — includes build-to-delete pruning and…

    141 GitHub stars~1.5k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Evals Context

    zgsm-ai/costrict

    Provides context about the CoStrict evals system structure in this monorepo.

    4.4k GitHub starsUsed in 1 repo~1.9k tokens
    AI & LLM EngineeringAuto-check passed
  • Fine-Tuning Expert

    Jeffallan/claude-skills

    Guides LLM fine-tuning with LoRA and QLoRA through Hugging Face PEFT, from dataset validation and training checks to adapter merging, quantization and deployment.

    12k GitHub starsUsed in 1 repo~1.7k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.

    11k GitHub starsUsed in 2 repos~1.6k tokens
    AI & LLM EngineeringAuto-check passed

More from evaleval/every_eval_ever

  • Eee Datastore PR Review

    evaleval/every_eval_ever

    Review and repair pull requests on the evaleval/EEEdatastore Hugging Face dataset.

    133 GitHub stars~2.5k tokensUpdated today
    Auto-check passed

Questions about Eee Dataset Conversion

What does Eee Dataset Conversion do?

Convert an evaluation dataset or leaderboard into the Every Eval Ever (EEE) schema — aggregate .json logs (eval.schema.json) and optional instance samples.jsonl sidecars…. Eee Dataset Conversion is an agent skill from evaleval/every_eval_ever.json).

When should I use Eee Dataset Conversion?

Eee Dataset Conversion fits situations like: asked to write an EEE adapter; add a dataset/leaderboard to the EEE datastore; map benchmark results into EEE; debug why an EEE record wont validate.

How do I install Eee Dataset Conversion in Claude Code?

Run `npx skills add evaleval/every_eval_ever --skill eee-dataset-conversion -a claude-code`. Or copy the skill folder (.agents/skills/eee-dataset-conversion in evaleval/every_eval_ever) into .claude/skills/eee-dataset-conversion in your project. Claude Code loads it when a task matches its description.

How do I install Eee Dataset Conversion in Codex?

Run `npx skills add evaleval/every_eval_ever --skill eee-dataset-conversion -a codex`. Or copy the skill folder (.agents/skills/eee-dataset-conversion in evaleval/every_eval_ever) into .agents/skills/eee-dataset-conversion in your project. Codex loads it when a task matches its description.

Can I use Eee Dataset Conversion in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add evaleval/every_eval_ever --skill eee-dataset-conversion -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/eee-dataset-conversion, .gemini/skills/eee-dataset-conversion, .github/skills/eee-dataset-conversion and .opencode/skills/eee-dataset-conversion in your project.

What does Eee Dataset Conversion need to run?

Going by SKILL.md and its folder, Eee Dataset Conversion needs Python and a shell for the scripts in its folder and the command-line tools its instructions call (uv). Our summary lists: Python 3; A Bash shell.

Does Eee Dataset Conversion access the network?

SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Eee Dataset Conversion safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Eee Dataset Conversion use?

Eee Dataset Conversion is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Eee Dataset Conversion use?

About 2.5k tokens (SKILL.md is roughly 9.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Eee Dataset Conversion?

Skills that share tags, products or a category with Eee Dataset Conversion: Azure AI Projects Python SDK (microsoft/skills, 3.1k stars), Spec Optimize (leo-kuang-ai/spec-first, 107 stars), Factory Learn (tikalk/adlc-team-skills, 141 stars) and Evals Context (zgsm-ai/costrict, 4.4k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Eee Dataset Conversion?

evaleval (a GitHub organization) maintains it in evaleval/every_eval_ever, which has 133 GitHub stars. The repository holds 2 skills in this directory. The repository was last updated on October 7, 2026.

Source: evaleval/every_eval_ever on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.