Agent skill

Flagevalmm Add Dataset

by flageval-baai in flageval-baai/FlagEvalMM

Integrate new evaluation datasets into FlagEvalMM as benchmark tasks.

No licenceAuto-check passedAI & LLM Engineering

Install Flagevalmm Add Dataset

skills CLI
$ npx skills add flageval-baai/FlagEvalMM --skill flagevalmm-add-dataset -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install flageval-baai/FlagEvalMM flagevalmm-add-dataset --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/flageval-baai/FlagEvalMM.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/flagevalmm-add-dataset .claude/skills/flagevalmm-add-dataset && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
flagevalmm-add-dataset
GitHub stars
108
Token cost
~3.2k tokens
SKILL.md length
728 words
Files
1
Skills in repo
1
Repo updated
First seen
Licence
None found

At a glance

Integrate new evaluation datasets into FlagEvalMM as benchmark tasks.

  • Works in 7 steps: Understand the Dataset → Write process.py → Write the Task Config → …
  • Adding a dataset from HuggingFace
  • SKILL.md covers Task Directory Structure, Workflow and Decision Tree
  • Reaches openrouter.ai

What it does

Flagevalmm Add Dataset is an agent skill from flageval-baai/FlagEvalMM. Integrate new evaluation datasets into FlagEvalMM as benchmark tasks. Use when adding a dataset from HuggingFace or other sources to FlagEvalMM, creating task configs, writing data processors, building custom evaluators, setting up prompt templates, or running evaluation benchmarks on VLMs. Trigger on: "add dataset to FlagEvalMM", "create a new task", "integrate benchmark", "evaluate model on [dataset]", "write process.py", "write evaluator", or any request involving the tasks/ directory of FlagEvalMM.

Its SKILL.md is about 3.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering Model hubs and datasets and Prompt engineering. It works with Hugging Face. The repository describes itself as: A Flexible Framework for Comprehensive Multimodal Model Evaluation.

When your agent uses it

  • Adding a dataset from HuggingFace
  • Other sources to FlagEvalMM
  • Creating task configs
  • Writing data processors

Example prompts

  • “add dataset to FlagEvalMM”
  • “create a new task”
  • “integrate benchmark”
  • “/flagevalmm-add-dataset”

Requirements

  • Python 3

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Understand the Dataset
  2. Write process.py
  3. Write the Task Config
  4. Choose Dataset Type
  5. Choose Evaluator
  6. Design the Prompt Template
  7. Test and Verify

What it can do on your machine

Read from SKILL.md and the folder at commit fef27ec. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python, bash and json).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • openrouter.ai

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Flagevalmm Add Dataset loads about 3.2k tokens when it runs. Until then it costs about 133 tokens; SKILL.md has 728 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~133
When it runs · the whole SKILL.md, loaded when a task matches
~3.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

Without a licence we can't republish the file, so here is its outline and opening line. It has 728 words (~3,225 tokens).

“Add new evaluation datasets to FlagEvalMM as benchmark tasks. This skill covers the full workflow: data processing, task configuration, evaluation logic, prompt design, and verification.”

— opening of SKILL.md by flageval-baai
name
flagevalmm-add-dataset

Read the full SKILL.md on GitHub

Files

Just SKILL.md in skills/flagevalmm-add-dataset of flageval-baai/FlagEvalMM.

Open the folder on GitHubat commit fef27ec

Compare with similar skills

Flagevalmm Add Dataset next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Flagevalmm Add Dataset compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Flagevalmm Add Dataset this skillflageval-baai/FlagEvalMM108—~3.2kAutomated safety check: PassNone
Hugging Face Datasetssickn33/agentic-awesome-skills47k2 repos~1.1kAutomated safety check: PassMIT
LLM Benchmarking with lm-evaluation-harnessOrchestra-Research/AI-Research-SKILLs13k8 repos~3kAutomated safety check: PassMIT
SageMaker Serving Image Selectionhuggingface/skills11k1 repos~4.6kAutomated safety check: PassApache-2.0
Hugging Face Local Model Evalshuggingface/skills11k2 repos~1.6kAutomated safety check: PassApache-2.0
Upload Post Imagehuggingface/blog3.5k—~1.1kAutomated safety check: PassNone

Similar skills

  • Hugging Face Datasets

    sickn33/agentic-awesome-skills

    Create and manage datasets on Hugging Face Hub. An agent skill from sickn33/agentic-awesome-skills.

    47k GitHub starsUsed in 2 repos~1.1k tokens
    AI & LLM EngineeringAuto-check passed
  • LLM Benchmarking with lm-evaluation-harness

    Orchestra-Research/AI-Research-SKILLs

    Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.

    13k GitHub starsUsed in 8 repos~3k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.

    11k GitHub starsUsed in 1 repo~4.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.

    11k GitHub starsUsed in 2 repos~1.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Upload Post Image

    huggingface/blog

    Official

    A skill your agent uses when adding or migrating non-thumbnail images for a Hugging Face Blog post.

    3.5k GitHub stars~1.1k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Esmfold2

    JimLiu/science-skills

    Biohub ESMFold2 / ESMFold2-Fast all-atom co-folding (Candido et al.

    228 GitHub starsUsed in 4 repos~2.5k tokens
    AI & LLM EngineeringAuto-check passed

Works with

Questions about Flagevalmm Add Dataset

What does Flagevalmm Add Dataset do?

Integrate new evaluation datasets into FlagEvalMM as benchmark tasks. Flagevalmm Add Dataset is an agent skill from flageval-baai/FlagEvalMM. Integrate new evaluation datasets into FlagEvalMM as benchmark tasks.

When should I use Flagevalmm Add Dataset?

Flagevalmm Add Dataset fits situations like: adding a dataset from HuggingFace; other sources to FlagEvalMM; creating task configs; writing data processors.

How do I install Flagevalmm Add Dataset in Claude Code?

Run `npx skills add flageval-baai/FlagEvalMM --skill flagevalmm-add-dataset -a claude-code`. Or copy the skill folder (skills/flagevalmm-add-dataset in flageval-baai/FlagEvalMM) into .claude/skills/flagevalmm-add-dataset in your project. Claude Code loads it when a task matches its description.

How do I install Flagevalmm Add Dataset in Codex?

Run `npx skills add flageval-baai/FlagEvalMM --skill flagevalmm-add-dataset -a codex`. Or copy the skill folder (skills/flagevalmm-add-dataset in flageval-baai/FlagEvalMM) into .agents/skills/flagevalmm-add-dataset in your project. Codex loads it when a task matches its description.

Can I use Flagevalmm Add Dataset in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add flageval-baai/FlagEvalMM --skill flagevalmm-add-dataset -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/flagevalmm-add-dataset, .gemini/skills/flagevalmm-add-dataset, .github/skills/flagevalmm-add-dataset and .opencode/skills/flagevalmm-add-dataset in your project.

What does Flagevalmm Add Dataset need to run?

SKILL.md names no scripts, command-line tools or credentials: Flagevalmm Add Dataset is instructions for the agent only. Our summary lists: Python 3.

Does Flagevalmm Add Dataset access the network?

SKILL.md names 1 domain. In commands or code: openrouter.ai; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Flagevalmm Add Dataset safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Flagevalmm Add Dataset use?

No licence was found for Flagevalmm Add Dataset or its repository. Without one, default copyright applies: ask the author before reusing or redistributing it.

How many tokens does Flagevalmm Add Dataset use?

About 3.2k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Flagevalmm Add Dataset?

Skills that share tags, products or a category with Flagevalmm Add Dataset: Hugging Face Datasets (sickn33/agentic-awesome-skills, 47k stars), LLM Benchmarking with lm-evaluation-harness (Orchestra-Research/AI-Research-SKILLs, 13k stars), SageMaker Serving Image Selection (huggingface/skills, 11k stars) and Hugging Face Local Model Evals (huggingface/skills, 11k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Flagevalmm Add Dataset?

flageval-baai (a GitHub organization) maintains it in flageval-baai/FlagEvalMM, which has 108 GitHub stars. The repository was last updated on April 21, 2026.

Source: flageval-baai/FlagEvalMM on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.