Agent skill

Eee Datastore PR Review

by evaleval in evaleval/every_eval_ever

Review and repair pull requests on the evaleval/EEEdatastore Hugging Face dataset.

MITAuto-check passedDevelopment

Install Eee Datastore PR Review

skills CLI
$ npx skills add evaleval/every_eval_ever --skill eee-datastore-pr-review -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install evaleval/every_eval_ever eee-datastore-pr-review --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/evaleval/every_eval_ever.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/eee-datastore-pr-review .claude/skills/eee-datastore-pr-review && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
eee-datastore-pr-review
GitHub stars
134
Token cost
~2.5k tokens
SKILL.md length
1,308 words
Files
3
Skills in repo
2
Repo updated
First seen
Licence
MIT

At a glance

Review and repair pull requests on the evaleval/EEEdatastore Hugging Face dataset.

  • Works in 8 steps: Establish the exact PR state → Reproduce the gate locally → Triage before editing → …
  • Given an EEEdatastore discussion
  • SKILL.md covers Operating contract, Load the live EEE rules, Workflow and Completion report
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Eee Datastore PR Review is an agent skill from evaleval/every_eval_ever. Review and repair pull requests on the evaleval/EEEdatastore Hugging Face dataset. Use when given an EEEdatastore discussion or PR URL, asked to run or reproduce /eee validate changed, resolve EEE validator errors or warnings, research model deploymenttype or modelavailability, edit the changed datastore records, rerun the bot, or prepare canonical-registry follow-ups.

Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files (for example `agents/openai.yaml` and `reference/model-deployment.md`).

It sits in Development, covering Pull requests and Model hubs and datasets. It works with Hugging Face. The repository describes itself as: Every Eval Ever is a shared schema and crowdsourced eval database. It defines a standardized metadata format for storing AI evaluation results — from leaderboard scrapes and… The licence is MIT.

When your agent uses it

  • Given an EEEdatastore discussion
  • Reproduce /eee validate changed
  • Resolve EEE validator errors
  • Research model deploymenttype

Example prompts

  • “/eee-datastore-pr-review”

Requirements

  • Python 3

Workflow steps

8 steps, taken from the step headings in SKILL.md.

  1. Establish the exact PR state
  2. Reproduce the gate locally
  3. Triage before editing
  4. Research ambiguous metadata
  5. Make the repair
  6. Verify the repaired head
  7. Update and monitor the existing PR
  8. Handle registry work without inventing a registry

What it can do on your machine

Read from SKILL.md and the folder at commit 1eb9d39. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Eee Datastore PR Review loads about 2.5k tokens when it runs. Until then it costs about 100 tokens; SKILL.md has 1,308 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~100
When it runs · the whole SKILL.md, loaded when a task matches
~2.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from evaleval/every_eval_ever at commit 1eb9d39, republished under its MIT licence (© evaleval). 1,308 words, ~2,495 tokens.

Download SKILL.mdSave it as .claude/skills/eee-datastore-pr-review/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
eee-datastore-pr-review
description
Review and repair pull requests on the evaleval/EEE_datastore Hugging Face dataset. Use when given an EEE_datastore discussion or PR URL, asked to run or reproduce `/eee validate changed`, resolve EEE validator errors or warnings, research model deployment_type or model_availability, edit the changed datastore records, rerun the bot, or prepare canonical-registry follow-ups.

Review and repair an EEE datastore PR

Produce the smallest source-backed change that makes the existing PR both validator-clean and semantically correct. Treat a green validator as necessary, not sufficient.

Operating contract

  • Treat the current checkout's schemas and REGISTERED_CHECKS as the local source of truth. Treat the newest bot result for the current PR head as the remote gate.
  • Report what the source establishes. Never invent metadata, clamp a score into its declared bounds, or change a value merely to silence the validator.
  • Make unknown a researched conclusion, not a default. Record which relevant surfaces were checked before retaining it.
  • Distinguish absent record metadata from unavailable source evidence. A missing or null model_info.additional_details object means the record needs investigation; it does not establish either axis as unknown.
  • Keep work on the supplied refs/pr/<number> ref. Do not open a replacement PR for another repair round.
  • If asked only to review, prepare a patch and findings without uploading or commenting. If asked to fix, update the supplied PR, trigger its validator, and iterate on that same ref.
  • Ask the operator before a policy decision: minting a new canonical id, changing a schema/validator rule, dropping non-trivial data, choosing an ambiguous metric or bound, or making another structural change. Do not hide such a choice in a data repair.

Load the live EEE rules

Before editing, read these sibling references:

  • ../eee-dataset-conversion/reference/datastore-gate.md
  • ../eee-dataset-conversion/reference/fields.md
  • ../eee-dataset-conversion/reference/datastore-submission.md
  • ../eee-dataset-conversion/reference/verification.md

Read reference/model-deployment.md whenever either model deployment axis is missing, stale, invalid, or suspicious. Read ../eee-dataset-conversion/reference/registry.md when an id is unresolved or a registry update is requested. Load the full eee-dataset-conversion skill when the repair also changes an adapter or regenerated output.

Re-read the allowed deployment values from every_eval_ever/validator/validation_core.py and the live schema. Existing records and old bot comments may use obsolete vocabularies.

Workflow

1. Establish the exact PR state
  1. Parse the dataset repo and discussion number from the supplied URL.
  2. Fetch the discussion details, commit history, current head, base ref, file diff, conflicts, and every validator comment. Prefer Hugging Face's API or huggingface_hub over scraping rendered HTML.
  3. Select only the newest completed bot run whose fingerprint or head matches the current PR. Older green runs describe older data or validator versions.
  4. Check out refs/pr/<number> in a dedicated datastore worktree or temporary clone. Preserve the contributor's branch and unrelated changes.
  5. Diff the PR head from its merge base with main. Inventory added, modified, renamed, and deleted paths; include aggregate/instance companions even if only one side appears in the diff.

Record the PR head commit and bot schema/compatibility version in the review notes. If the bot and local schema differ, label their disagreement as version skew and investigate it explicitly.

2. Reproduce the gate locally

Run the current EEE CLI against changed .json and .jsonl files at their final data/<collection>/<developer>/<model>/... paths. Pass files or a quoted glob, never a directory. Include companion files required by semantic validation.

Use:

text
uv run python -m every_eval_ever validate <changed files>
uv run python -m every_eval_ever.check_duplicate_entries <relevant files>

Capture the full output and exit status. Do not rely on Pydantic model construction or validate_file() alone; those can omit semantic checks. If current main and the deployed bot disagree, reproduce both versions when practical and fix toward the current schema without silently degrading data for an old bot.

3. Triage before editing

Group findings by root cause rather than by file. For each group, record:

  • affected paths and exact model/result identities;
  • local and bot messages;
  • whether the issue is mechanical, schema-semantic, or content-semantic;
  • the source evidence needed for a correct fix;
  • proposed change and confidence.

Inspect content even when the validator omits it. At minimum check suspicious zeroes, score scale and bounds, metric identity, source_data, duplicate overall/subtask aggregates, stable evaluation_id, model identity, answer leakage, and companion pairing. An out-of-range score requires finding the source scale or source value; do not cap, clamp, or round it into validity.

Inspect the raw JSON before constructing an EvaluationLog. The model layer may auto-fill absent deployment keys with unknown, hiding whether the contributor actually supplied additional_details, supplied only one axis, or supplied neither.

4. Research ambiguous metadata

For deployment warnings, apply reference/model-deployment.md to each exact model variant and evaluation run. Determine the two axes independently. Do not infer one from the other, from the developer folder, or from a provider-wide rule.

Search all relevant primary surfaces before choosing unknown: record payload and run config, generating adapter, pinned model card, evaluator methodology, paper and appendix, source repository, and official API/release documentation. Use current web research where facts may have changed, but pin the evidence revision or date relevant to the submitted evaluation.

Batch models only after proving that they share the same evidence. Keep an evidence table with raw model label, canonical model id, both decisions, source URL/revision, and confidence.

Show full SKILL.md (531 more words)Show less
5. Make the repair
  • Edit only files implicated by a finding. Avoid mass reformatting unrelated data.
  • Preserve UUID filenames and stable evaluation identities unless identity itself is the defect.
  • When model_info.additional_details is absent or null, create the object only after researching both axes. When it already exists, merge the researched keys without discarding unrelated source metadata.
  • Keep additional_details values as strings. Add concise evidence/provenance there when the source has no typed home and the decision would otherwise be opaque.
  • If generated records are wrong, fix or prepare the generating adapter in the code repo as well; otherwise the next refresh will restore the defect. Keep adapter code out of the datastore PR and cross-link its separate PR.
  • Review the resulting diff for accidental deletion, unrelated churn, and a mechanical replacement applied to semantically different models.
6. Verify the repaired head

Rerun the local validator and duplicate checker, then repeat the content spot-check. Require every changed file and companion to pass. Review warnings even if the command or bot says “Ready to Merge.”

Compare the final changed-path inventory with the initial inventory. Explain every new path, deletion, identity change, or source-value change in the decision log.

7. Update and monitor the existing PR

When the task authorizes a fix, upload exact add/delete operations to the existing refs/pr/<number> with huggingface_hub.HfApi.create_commit; set the current PR head as parent_commit so concurrent updates fail instead of being overwritten. Never set create_pr=True for a repair round.

After the commit lands:

  1. Comment /eee validate changed on the same discussion with HfApi.comment_discussion.
  2. Monitor until a completed run matches the new head/fingerprint.
  3. Re-read every error and warning, repair locally, and repeat on the same ref.
  4. Stop only when both the current local CLI and matching bot run are clean, or when a genuine policy/ambiguity/auth/conflict blocker needs the operator.

Do not post claims or comments on the contributor's behalf during a review-only task.

8. Handle registry work without inventing a registry

Resolve model, benchmark, metric, harness, and organization ids against the registry when a resolver or registry checkout exists. Search existing canonicals and aliases before proposing anything new.

If the registry repository and its contribution workflow are available:

  1. Read its AGENTS.md, CONTRIBUTING.md, and registry skill.
  2. Add an alias to an existing canonical when evidence supports it.
  3. Ask the operator before deliberately creating a new canonical.
  4. Validate in that repo and open a separate registry PR; cross-link it with the data and adapter PRs.

If the registry is unavailable or not yet implemented, do not invent its file format. Emit a registry-candidate table in the review report with entity type, raw value, candidate canonical, evidence, confidence, and whether the candidate is an alias or a new entity. Leave the datastore value source-faithful and mark resolution status explicitly.

Completion report

Return:

  • PR URL, starting head, final head, and matching bot run/version;
  • files changed, grouped by root cause;
  • local validation and duplicate-check results;
  • content spot-checks performed;
  • deployment/availability evidence table, including researched unknown values;
  • registry and adapter follow-ups with cross-links or candidate tables;
  • decision log and any unresolved blocker.

Do not call a PR complete merely because all files parse or the bot prints “Ready to Merge.”

© evaleval, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in .agents/skills/eee-datastore-pr-review of evaleval/every_eval_ever.

  • SKILL.md
  • agents/openai.yaml
  • reference/model-deployment.md

Open the folder on GitHubat commit 1eb9d39

Compare with similar skills

Eee Datastore PR Review next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Eee Datastore PR Review compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Eee Datastore PR Review this skillevaleval/every_eval_ever134—~2.5kAutomated safety check: PassMIT
Codex Commit AgentJulian-adv/OpenMMO1.8k—~1.3kAutomated safety check: PassCustom licence
Hugging Face API Tool Builderhuggingface/skills11k2 repos~1.5kAutomated safety check: PassApache-2.0
Quark Torch Shrink Modelamd/Quark181—~980Automated safety check: PassMIT
Comfy CLIsundial-org/awesome-openclaw-skills663—~1.5kAutomated safety check: PassNone
SageMaker Serving Image Selectionhuggingface/skills11k1 repos~4.6kAutomated safety check: PassApache-2.0

Similar skills

  • Codex Commit Agent

    Julian-adv/OpenMMO

    Run the repository commit workflow when the user explicitly asks to commit changes or create a save-point commit.

    1.8k GitHub stars~1.3k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Official

    Builds reusable command line scripts that fetch, enrich or process data from the Hugging Face API, aimed at chained, repeated or automated tasks.

    11k GitHub starsUsed in 2 repos~1.5k tokens
    AI & LLM EngineeringAuto-check passed
  • Shrink a HuggingFace safetensors model to 1 hidden layer for fast debugging without loading the full model into memory.

    181 GitHub stars~980 tokensUpdated 12 days ago
    AI & LLM EngineeringAuto-check passed
  • Comfy CLI

    sundial-org/awesome-openclaw-skills

    Install, manage, and run ComfyUI instances. An agent skill from sundial-org/awesome-openclaw-skills.

    663 GitHub stars~1.5k tokensUpdated 7 mo ago
    AI & LLM EngineeringAuto-check passed
  • Official

    Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.

    11k GitHub starsUsed in 1 repo~4.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.

    11k GitHub starsUsed in 2 repos~1.6k tokens
    AI & LLM EngineeringAuto-check passed

More from evaleval/every_eval_ever

  • Eee Dataset Conversion

    evaleval/every_eval_ever

    Convert an evaluation dataset or leaderboard into the Every Eval Ever (EEE) schema — aggregate .json logs (eval.schema.json) and optional instance samples.jsonl sidecars…

    134 GitHub stars~2.5k tokensUpdated 2 days ago
    Auto-check passed

Works with

Questions about Eee Datastore PR Review

What does Eee Datastore PR Review do?

Review and repair pull requests on the evaleval/EEEdatastore Hugging Face dataset. Eee Datastore PR Review is an agent skill from evaleval/every_eval_ever. Review and repair pull requests on the evaleval/EEEdatastore Hugging Face dataset.

When should I use Eee Datastore PR Review?

Eee Datastore PR Review fits situations like: given an EEEdatastore discussion; reproduce /eee validate changed; resolve EEE validator errors; research model deploymenttype.

How do I install Eee Datastore PR Review in Claude Code?

Run `npx skills add evaleval/every_eval_ever --skill eee-datastore-pr-review -a claude-code`. Or copy the skill folder (.agents/skills/eee-datastore-pr-review in evaleval/every_eval_ever) into .claude/skills/eee-datastore-pr-review in your project. Claude Code loads it when a task matches its description.

How do I install Eee Datastore PR Review in Codex?

Run `npx skills add evaleval/every_eval_ever --skill eee-datastore-pr-review -a codex`. Or copy the skill folder (.agents/skills/eee-datastore-pr-review in evaleval/every_eval_ever) into .agents/skills/eee-datastore-pr-review in your project. Codex loads it when a task matches its description.

Can I use Eee Datastore PR Review in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add evaleval/every_eval_ever --skill eee-datastore-pr-review -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/eee-datastore-pr-review, .gemini/skills/eee-datastore-pr-review, .github/skills/eee-datastore-pr-review and .opencode/skills/eee-datastore-pr-review in your project.

What does Eee Datastore PR Review need to run?

SKILL.md names no scripts, command-line tools or credentials: Eee Datastore PR Review is instructions for the agent only. Our summary lists: Python 3.

Does Eee Datastore PR Review access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Eee Datastore PR Review safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Eee Datastore PR Review use?

Eee Datastore PR Review is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Eee Datastore PR Review use?

About 2.5k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Eee Datastore PR Review?

Skills that share tags, products or a category with Eee Datastore PR Review: Codex Commit Agent (Julian-adv/OpenMMO, 1.8k stars), Hugging Face API Tool Builder (huggingface/skills, 11k stars), Quark Torch Shrink Model (amd/Quark, 181 stars) and Comfy CLI (sundial-org/awesome-openclaw-skills, 663 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Eee Datastore PR Review?

evaleval (a GitHub organization) maintains it in evaleval/every_eval_ever, which has 134 GitHub stars. The repository holds 2 skills in this directory. The repository was last updated on October 7, 2026.

Source: evaleval/every_eval_ever on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.