Official agent skill

Dataset Evaluation

by awslabs in awslabs/agent-plugins

Validates dataset formatting and quality for SageMaker model fine-tuning (SFT, DPO, or RLVR).

OfficialApache-2.0Auto-check passedAI & LLM Engineering

Install Dataset Evaluation

skills CLI
$ npx skills add awslabs/agent-plugins --skill dataset-evaluation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install awslabs/agent-plugins dataset-evaluation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/awslabs/agent-plugins.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/sagemaker-ai/skills/dataset-evaluation .claude/skills/dataset-evaluation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
dataset-evaluation
GitHub stars
915
Used in
2 other repos
Token cost
~1.3k tokens
SKILL.md length
567 words
Files
4 (incl. scripts, references)
Skills in repo
33
Repo updated
First seen
Licence
Apache-2.0

At a glance

Validates dataset formatting and quality for SageMaker model fine-tuning (SFT, DPO, or RLVR).

  • Works in 4 steps: Locate Dataset → Determine strategy and model → Check File Formatting: Run the tool… → …
  • The user says is my dataset okay
  • SKILL.md covers Prerequisites, Workflow, Messages to the User and Script Details, plus 1 more section
  • Runs Python scripts from its folder; calls python

What it does

Dataset Evaluation is an agent skill from awslabs/agent-plugins, published by the product's own GitHub organization. Validates dataset formatting and quality for SageMaker model fine-tuning (SFT, DPO, or RLVR). Use when the user says "is my dataset okay", "evaluate my data", "check my training data", "I have my own data", or before starting any fine-tuning job. Detects file format, checks schema compliance against the selected model and technique, and reports whether the data is ready for training or evaluation.

Its SKILL.md is about 1.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including scripts and reference files (for example `references/custom-scorer-evaluation-dataset-formats.md`, `references/strategy_data_requirements.md` and `scripts/format_detector.py`).

It sits in AI & LLM Engineering, covering Fine-tuning. It works with Amazon SageMaker. The repository describes itself as: Agent Plugins for AWS equip AI coding agents with the skills to help you architect, deploy, and operate on AWS. The licence is Apache-2.0.

When your agent uses it

  • The user says is my dataset okay
  • Evaluate my data
  • Check my training data
  • I have my own data

Example prompts

  • “is my dataset okay”
  • “evaluate my data”
  • “check my training data”
  • “/dataset-evaluation”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Locate Dataset
  2. Determine strategy and model
  3. Check File Formatting: Run the tool format_detector.py to make sure the file conforms to formatting requirements.
  4. Summarize Results: Tell the user if their data is ready

What it can do on your machine

Read from SKILL.md and the folder at commit da51970. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • docs.aws.amazon.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Dataset Evaluation loads about 1.3k tokens when it runs, and up to ~3.5k if it reads all its reference files. Until then it costs about 105 tokens; SKILL.md has 567 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~105
When it runs · the whole SKILL.md, loaded when a task matches
~1.3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~3.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from awslabs/agent-plugins at commit da51970, republished under its Apache-2.0 licence (© awslabs). 567 words, ~1,260 tokens.

Download SKILL.mdSave it as .claude/skills/dataset-evaluation/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
dataset-evaluation
description
Validates dataset formatting and quality for SageMaker model fine-tuning (SFT, DPO, or RLVR). Use when the user says "is my dataset okay", "evaluate my data", "check my training data", "I have my own data", or before starting any fine-tuning job. Detects file format, checks schema compliance against the selected model and technique, and reports whether the data is ready for training or evaluation.
metadata.version
1.0.0

Workflow Instruction

Follow the workflow shown below. Locate the dataset, check the file type, and resolve any issues with missing files or wrong file types. Determine the fine-tuning model and fine-tuning strategy. Run the appropriate validation based on the model family. Summarize the results: is the dataset ready for fine-tuning?

Prerequisites

  • The SDK environment has been verified (SDK version, region, execution role). If not done, activate the sdk-getting-started skill first.

Workflow

  1. Locate Dataset:

    • The full path may be a local file path, or an S3 URI
    • Resolve the full path to the dataset file, make sure read permissions are available, and help the user if the file is not found
  2. Determine strategy and model:

    • File formatting depends on the currently selected fine-tuning strategy and fine-tuning base model.
    • If the strategy and model are already known from the conversation context (e.g., selected via the model-selection and finetuning-technique skills), use them.
    • If not available in context, activate the model-selection and/or finetuning-technique skills to determine them before proceeding.
    • Exception: If the user is validating an evaluation dataset (not a training dataset), neither model nor technique is required — the format detector can validate eval format (query/response structure) independently. Do not block on model-selection or finetuning-technique for eval dataset validation.
  3. Check File Formatting: Run the tool format_detector.py to make sure the file conforms to formatting requirements.

    • Send the full path directly to the format_detector script as an argument
    • Do not send the model and strategy as arguments
    • Do not download data from S3
    • Do not make local copies of data
  4. Summarize Results: Tell the user if their data is ready

    • Examine the output of format_detector and compare to the known strategy and model
    • Important: training datasets and evaluation datasets have different format requirements.
      • Training datasets must match the fine-tuning strategy format per references/strategy_data_requirements.md
      • Evaluation datasets (for model evaluation) must match one of the SageMaker evaluation dataset formats.
      • Custom Scorer evaluation datasets have scorer-specific requirements. If the dataset is intended for Custom Scorer evaluation (Prime Math, Prime Code, or Custom Lambda), read references/custom-scorer-evaluation-dataset-formats.md and validate against the scorer-specific schema. The scorer type should be known from conversation context (determined in the model-evaluation skill).
    • Report back to the user if their current dataset is valid for its intended purpose
    • Warn the user if their dataset is valid, but for a different strategy or model
    • Warn the user if their dataset is not valid for any strategy/model pair
    • If the user plans to finetune a model with the evaluated dataset, it needs to be uploaded to an S3 bucket in the same region as the planned training job (usually the default region). Warn the user if this is NOT the case.
    • If the dataset is NOT in the necessary format, recommend transforming it using the dataset-transformation skill, wait for user confirmation, and update the plan based on their response
Show full SKILL.md (92 more words)Show less

Messages to the User

  • Introduction: "This skill checks the structure of your dataset for model fine-tuning."
  • File types: This skill applies to files that are formatted according to the Amazon SageMaker AI Developer Guide

Resources

  • scripts/format_detector.py is self-contained format validation script that can be run independently
  • model-selection and finetuning-technique skills should have already determined the base model and fine-tuning strategy
  • references/strategy_data_requirements.md contains data format requirements per strategy

Script Details

  • scripts/format_detector.py is self-contained format validation script that can be run independently:
bash
# With the file path argument identified in workflow step 1
python scripts/format_detector.py local_path/to/dataset

References

  • scripts/format_detector.py — Self-contained format validation script
  • references/strategy_data_requirements.md — Data format requirements per strategy

© awslabs, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (scripts, references) in plugins/sagemaker-ai/skills/dataset-evaluation of awslabs/agent-plugins.

  • SKILL.md
  • references/custom-scorer-evaluation-dataset-formats.md
  • references/strategy_data_requirements.md
  • scripts/format_detector.py

Open the folder on GitHubat commit da51970

Used in 2 other repositories

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in awslabs/agent-plugins, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Dataset Evaluation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Dataset Evaluation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Dataset Evaluation this skillawslabs/agent-plugins9152 repos~1.3kAutomated safety check: PassApache-2.0
AWS AI MLaws/agent-toolkit-for-aws2.8k—~1.7kAutomated safety check: PassApache-2.0
Peft Fine TuningOrchestra-Research/AI-Research-SKILLs13k9 repos~3.1kAutomated safety check: PassMIT
SageMaker Serving Image Selectionhuggingface/skills11k1 repos~4.6kAutomated safety check: PassApache-2.0
Hugging Face LLM Trainerhuggingface/skills11k3 repos~7.2kAutomated safety check: PassApache-2.0
Sentence-Transformers Training Routerhuggingface/skills11k1 repos~2.6kAutomated safety check: PassApache-2.0

Similar skills

  • AWS AI ML

    aws/agent-toolkit-for-aws

    Official

    Selects, deploys, and customizes AI models on Amazon SageMaker.

    2.8k GitHub stars~1.7k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Peft Fine Tuning

    Orchestra-Research/AI-Research-SKILLs

    Parameter-efficient fine-tuning for LLMs using LoRA, QLoRA, and 25+ methods.

    13k GitHub starsUsed in 9 repos~3.1k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.

    11k GitHub starsUsed in 1 repo~4.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Hugging Face LLM Trainer

    huggingface/skills

    Official

    Trains or fine-tunes language and vision models with TRL or Unsloth on Hugging Face Jobs cloud GPUs, then converts the results to GGUF.

    11k GitHub starsUsed in 3 repos~7.2k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Routes a sentence-transformers training task to the right model type and required reference docs and example scripts, covering bi-encoders, rerankers, sparse and multi-vector models.

    11k GitHub starsUsed in 1 repo~2.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Train Rl

    OpenPipe/ART

    RL training reference for the ART framework. An agent skill from OpenPipe/ART.

    11k GitHub stars~2.4k tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from awslabs/agent-plugins

All 33 skills in this repo
  • Dataset Transformation

    awslabs/agent-plugins

    Official

    Generates code that transforms datasets between ML schemas for model training or evaluation.

    915 GitHub starsUsed in 2 repos~3.5k tokens
    Auto-check passed
  • Finetuning Technique

    awslabs/agent-plugins

    Official

    Selects a fine-tuning technique (SFT, DPO, RLVR, or RLAIF) for the user's use case and validates it against the selected model's available recipes.

    915 GitHub starsUsed in 1 repo~604 tokens
    Auto-check passed
  • Hyperpod Issue Report

    awslabs/agent-plugins

    Official

    Generate comprehensive issue reports from HyperPod clusters (EKS and Slurm) by collecting diagnostic logs and configurations for troubleshooting and AWS Support cases.

    915 GitHub starsUsed in 1 repo~890 tokens
    Auto-check passed
  • Hyperpod Performance Debugger

    awslabs/agent-plugins

    Official

    Diagnose performance issues on Amazon SageMaker HyperPod clusters — uneven NCCL bandwidth across nodes and poor filesystem throughput.

    915 GitHub starsUsed in 1 repo~4.1k tokens
    Auto-check passed
  • Hyperpod Ssm

    awslabs/agent-plugins

    Official

    Remote command execution and file transfer on SageMaker HyperPod cluster nodes via AWS Systems Manager (SSM).

    915 GitHub starsUsed in 1 repo~1.3k tokens
    Auto-check: notes
  • AWS Lambda Managed Instances

    awslabs/agent-plugins

    Official

    Evaluate, configure, and migrate workloads to AWS Lambda Managed Instances (LMI).

    915 GitHub stars~4k tokensUpdated 2 days ago
    Auto-check passed

Questions about Dataset Evaluation

What does Dataset Evaluation do?

Validates dataset formatting and quality for SageMaker model fine-tuning (SFT, DPO, or RLVR). Dataset Evaluation is an agent skill from awslabs/agent-plugins, published by the product's own GitHub organization. Validates dataset formatting and quality for SageMaker model fine-tuning (SFT, DPO, or RLVR).

When should I use Dataset Evaluation?

Dataset Evaluation fits situations like: the user says is my dataset okay; evaluate my data; check my training data; I have my own data.

How do I install Dataset Evaluation in Claude Code?

Run `npx skills add awslabs/agent-plugins --skill dataset-evaluation -a claude-code`. Or copy the skill folder (plugins/sagemaker-ai/skills/dataset-evaluation in awslabs/agent-plugins) into .claude/skills/dataset-evaluation in your project. Claude Code loads it when a task matches its description.

How do I install Dataset Evaluation in Codex?

Run `npx skills add awslabs/agent-plugins --skill dataset-evaluation -a codex`. Or copy the skill folder (plugins/sagemaker-ai/skills/dataset-evaluation in awslabs/agent-plugins) into .agents/skills/dataset-evaluation in your project. Codex loads it when a task matches its description.

Can I use Dataset Evaluation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add awslabs/agent-plugins --skill dataset-evaluation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/dataset-evaluation, .gemini/skills/dataset-evaluation, .github/skills/dataset-evaluation and .opencode/skills/dataset-evaluation in your project.

What does Dataset Evaluation need to run?

Going by SKILL.md and its folder, Dataset Evaluation needs Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python 3.

Does Dataset Evaluation access the network?

SKILL.md names 1 domain. As links in the text: docs.aws.amazon.com. This is read from the text; nothing was executed.

Is Dataset Evaluation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Dataset Evaluation use?

Dataset Evaluation is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Dataset Evaluation use?

About 1.3k tokens (SKILL.md is roughly 5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.3k tokens, read only when the agent opens those files.

What are the alternatives to Dataset Evaluation?

Skills that share tags, products or a category with Dataset Evaluation: AWS AI ML (aws/agent-toolkit-for-aws, 2.8k stars), Peft Fine Tuning (Orchestra-Research/AI-Research-SKILLs, 13k stars), SageMaker Serving Image Selection (huggingface/skills, 11k stars) and Hugging Face LLM Trainer (huggingface/skills, 11k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Dataset Evaluation?

awslabs (a GitHub organization, an official publisher) maintains it in awslabs/agent-plugins, which has 915 GitHub stars. The repository holds 33 skills in this directory. The repository was last updated on October 5, 2026.

Source: awslabs/agent-plugins on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.