SageMaker Serving Image Selection
huggingface/skills
Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.
Generates code that transforms datasets between ML schemas for model training or evaluation.
$ npx skills add awslabs/agent-plugins --skill dataset-transformation -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install awslabs/agent-plugins dataset-transformation --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/awslabs/agent-plugins.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/sagemaker-ai/skills/dataset-transformation .claude/skills/dataset-transformation && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "dataset-transformation" agent skill from https://github.com/awslabs/agent-plugins/tree/main/plugins/sagemaker-ai/skills/dataset-transformation into .claude/skills/dataset-transformation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "dataset-transformation", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/awslabs/agent-plugins/tree/main/plugins/sagemaker-ai/skills/dataset-transformationType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add awslabs/agent-plugins --skill dataset-transformation -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install awslabs/agent-plugins dataset-transformation --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/awslabs/agent-plugins.git skills-src && mkdir -p .agents/skills && cp -r skills-src/plugins/sagemaker-ai/skills/dataset-transformation .agents/skills/dataset-transformation && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "dataset-transformation" agent skill from https://github.com/awslabs/agent-plugins/tree/main/plugins/sagemaker-ai/skills/dataset-transformation into .agents/skills/dataset-transformation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "dataset-transformation", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add awslabs/agent-plugins --skill dataset-transformation -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install awslabs/agent-plugins dataset-transformation --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/awslabs/agent-plugins.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/plugins/sagemaker-ai/skills/dataset-transformation .cursor/skills/dataset-transformation && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "dataset-transformation" agent skill from https://github.com/awslabs/agent-plugins/tree/main/plugins/sagemaker-ai/skills/dataset-transformation into .cursor/skills/dataset-transformation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "dataset-transformation", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/awslabs/agent-plugins.git --path plugins/sagemaker-ai/skills/dataset-transformation--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add awslabs/agent-plugins --skill dataset-transformation -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install awslabs/agent-plugins dataset-transformation --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/awslabs/agent-plugins.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/plugins/sagemaker-ai/skills/dataset-transformation .gemini/skills/dataset-transformation && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "dataset-transformation" agent skill from https://github.com/awslabs/agent-plugins/tree/main/plugins/sagemaker-ai/skills/dataset-transformation into .gemini/skills/dataset-transformation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "dataset-transformation", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install awslabs/agent-plugins dataset-transformationInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add awslabs/agent-plugins --skill dataset-transformation -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/awslabs/agent-plugins.git skills-src && mkdir -p .github/skills && cp -r skills-src/plugins/sagemaker-ai/skills/dataset-transformation .github/skills/dataset-transformation && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "dataset-transformation" agent skill from https://github.com/awslabs/agent-plugins/tree/main/plugins/sagemaker-ai/skills/dataset-transformation into .github/skills/dataset-transformation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "dataset-transformation", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add awslabs/agent-plugins --skill dataset-transformation -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install awslabs/agent-plugins dataset-transformation --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/awslabs/agent-plugins.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/plugins/sagemaker-ai/skills/dataset-transformation .opencode/skills/dataset-transformation && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "dataset-transformation" agent skill from https://github.com/awslabs/agent-plugins/tree/main/plugins/sagemaker-ai/skills/dataset-transformation into .opencode/skills/dataset-transformation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "dataset-transformation", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
dataset-transformationGenerates code that transforms datasets between ML schemas for model training or evaluation.
Dataset Transformation is an agent skill from awslabs/agent-plugins, published by the product's own GitHub organization. Generates code that transforms datasets between ML schemas for model training or evaluation. Use when the user says "transform", "convert", "reformat", "change the format", or when a dataset's schema needs to change to match the target format — always use this skill for format changes rather than writing inline transformation code. Supports OpenAI chat, SageMaker SFT/DPO/RLVR/RLAIF, HuggingFace preference, Bedrock Nova, VERL, and custom JSONL formats from local files or S3.
Its SKILL.md is about 3.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 8 other files, including scripts and reference files (for example `code_templates/transformation.py`, `references/code_output_guide.md` and `references/dataset_transformation_code.md`).
It sits in AI & LLM Engineering, covering Fine-tuning, File uploads and storage and Model hubs and datasets. It works with Amazon SageMaker, Hugging Face and OpenAI. The repository describes itself as: Agent Plugins for AWS equip AI coding agents with the skills to help you architect, deploy, and operate on AWS. The licence is Apache-2.0.
11 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit da51970. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
python3From the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
docs.aws.amazon.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Dataset Transformation loads about 3.5k tokens when it runs, and up to ~6.3k if it reads all its reference files. Until then it costs about 125 tokens; SKILL.md has 2,011 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from awslabs/agent-plugins at commit da51970, republished under its Apache-2.0 licence (© awslabs). 2,011 words, ~3,525 tokens.
.claude/skills/dataset-transformation/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.Transforms a data set provided by the user into their desired format.
sdk-getting-started skill first..jsonl (JSON Lines — one JSON object per line).This skill supports two transformation purposes — training data and evaluation data — each with its own format resolution path. The purpose is determined in Step 1 of the workflow.
Resolve the target format using the reference file ../dataset-evaluation/references/strategy_data_requirements.md. When the transformation is for model training, the required format depends on both the model type (Open Weights like Llama/Qwen vs Nova) and the finetuning technique (SFT, DPO, RLVR, RLAIF) — make sure to match on both dimensions. If either the model type or technique is not yet known, ask the user before resolving the format.
When the transformation is for model evaluation, resolve the target format using this order:
references/sagemaker_dataset_formats.md. Inform the user that the format schemas are from an offline copy and may be outdated.Use whichever source you successfully access as the source of truth for the target format. Do not rely on memorized schemas.
Your first response should determine whether this transformation is for model training or model evaluation. If the context already makes this clear (e.g., the user said "I need to prep my training data" or "I need to format my eval dataset"), confirm your understanding and move on. Otherwise, ask:
"Is this dataset transformation for model training or model evaluation? This helps me look up the right target format for you."
Remember this choice — it determines how the target format is resolved in Step 3.
⏸ Wait for user.
Acknowledge the user's request and state what this skill can do:
"I can help you transform your dataset's format! Here's my plan: I will first need to understand the format of your dataset and the transformation requirements. Once I have that, I will generate a dataset transformation function that we can refine together. After the dataset transformation function is refined to your liking, I will perform the transformation task and upload it to your desired location! Does this sound good?"
⏸ Wait for user.
For this step, you need to know: what dataset format the user would like to transform their dataset from and what dataset format they would like to transform it in to. If you know this already, skip this step. If not, ask the user:
"What's the dataset format you would like to transform it into?"
Resolve the target format based on the purpose determined in Step 1:
"I've found a SageMaker dataset format: {sagemaker-dataset-format-name} with schema: {sagemaker-dataset-format-schema}. Is this what you were referring to?"
If the user describes a custom format not listed in the reference doc, ask them to provide a sample record of the desired output format.
⏸ Wait for user.
For this step, you need: the location of the user's dataset. If you know this already, skip this step. If not, ask the user:
"Where can I find your dataset? Either a local directory or S3 location works!"
⏸ Wait for user.
Read 1–2 sample records from the user's dataset and show them so the user can confirm the source schema. Do not run format detection — that is handled by the planning skill before this skill is invoked.
Do not show a side-by-side mapping to the target format here — the detailed mapping will be handled in Step 7 when generating the transformation function.
⏸ Wait for user.
For this step, you need: to understand where to output the transformed dataset to. It could be an S3 URI or local directory If you already know where the dataset is supposed to be output to, skip this step. If not, ask the user:
"Where should I output your transformed dataset to? Either a local directory or S3 location works!"
If the user provides a directory (not a full file path), construct the output filename using the pattern {original_name}_{target_format}.jsonl (e.g., gen_qa_100k_openai.jsonl).
⏸ Wait for user.
For this step, you need: to generate a python function that transforms the dataset from the format in Step 5 to the format in Step 3
Read the reference guide at references/dataset_transformation_code.md and follow its skeleton exactly when generating the transformation function.
The python function should be in the form of:
def transform_dataset(df: pd.DataFrame) -> pd.DataFrame:The <project-dir> is the project directory established by the directory-management skill (e.g., dpo-to-rlvr-conversion).
In notebook mode, add a %%writefile <project-dir>/scripts/transform_fn.py code cell AND write the file to disk for testing. In script mode, write the file to disk directly.
Continue iterating with the user's feedback — update the code in place on each revision rather than showing code inline.
If sample data was collected in Step 5, test the function against the sample records:
/tmp/test_input.jsonl), then run:
python3 -c "import sys; sys.path.insert(0, '<project-dir>/scripts'); from transform_fn import transform_dataset; import pandas as pd; df = pd.read_json('/tmp/test_input.jsonl', lines=True); result = transform_dataset(df); print(result.to_json(orient='records', lines=True))"If no sample data, present the function for review and refinement.
⏸ Wait for user.
If no project directory exists, activate the directory-management skill to set one up.
⏸ Wait for user.
Before writing the code, read:
references/code_output_guide.md (output format rules)code_templates/transformation.py (cell structure and skeleton code)The template uses # Cell N: Label markers — each marker starts a new section. Cell 2 (Transformation Function) is dynamically generated from Step 7; all other cells follow the template skeleton.
Generate the execution logic following the code output guide.
%%writefile <project-dir>/scripts/<script_name>.py code cell AND write the file to disk. In script mode, write the file to disk directly.transform_dataset from transform_fn.Read the reference guide at references/dataset_transformation_code.md and follow its execution script skeleton exactly.
If sample data was collected in Step 5, test the full pipeline:
/tmp/test_input.jsonl).python3 <project-dir>/scripts/<script_name> --input /tmp/test_input.jsonl --output /tmp/test_output.jsonlIf no sample data, present the notebook for review and refinement.
⏸ Wait for user.
Check the size of the input dataset:
head-object (S3 service) with the bucket and key to get ContentLength.Decision criteria:
Inform the user of the recommendation and get their approval:
If local:
"Your dataset is {size} MB — since it's under 50 MB, I'd recommend running the transformation locally. Would you like to proceed with local execution, or would you prefer a SageMaker Processing Job instead?"
If SageMaker Processing Job:
"Your dataset is {size} MB — since it's over 50 MB, I'd recommend running this as a SageMaker Processing Job for better performance. Would you like to proceed with a SageMaker Processing Job, or would you prefer to run it locally instead?"
Do not execute until the user approves. If the user rejects the recommendation, switch to the alternative and get their explicit approval before proceeding.
⏸ Wait for user.
After user confirms, add an execution cell to the notebook. Do NOT run the transformation directly (no bash, no inline python). If notebook execution tools (run_cell) are available, offer to run the cells. Otherwise, generate the cell for the user to execute themselves:
If local execution:
.py files already on disk (written by the agent during Steps 7 and 9): import transform_dataset from transform_fn, load the dataset, transform, and save output. Scripts are located in <project-dir>/scripts/.If SageMaker Processing Job:
processor.run(wait=True, logs=True) to block the cell and stream logs until the job completes. See scripts/transformation_tools.py for reference implementation details.Important: The agent must NOT execute the transformation directly via bash or inline python. If run_cell is available, use it to run the notebook cells. Otherwise, the cells are for the user to review and run. Only sample data (from Steps 7 and 9) should be transformed by the agent for validation purposes.
If
run_cellis available: "I've added the execution cell to the notebook. Would you like me to run it?" Otherwise: "I've added the execution cell to the notebook. You can run it to transform the full dataset. Would you like to review the notebook before running it?"
⏸ Wait for user.
For this step, you need: to verify the output looks correct and confirm with the user.
⏸ Wait for user to confirm.
© awslabs, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 5 other files (scripts, references) in plugins/sagemaker-ai/skills/dataset-transformation of awslabs/agent-plugins.
Open the folder on GitHubat commit da51970
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in awslabs/agent-plugins, which our catalogue first saw on October 7, 2026.
Dataset Transformation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Dataset Transformation this skillawslabs/agent-plugins | 915 | 1 repos | ~3.5k | Automated safety check: Pass | Apache-2.0 | |
| SageMaker Serving Image Selectionhuggingface/skills | 11k | 1 repos | ~4.6k | Automated safety check: Pass | Apache-2.0 | |
| Hugging Face LLM Trainerhuggingface/skills | 11k | 1 repos | ~7.2k | Automated safety check: Pass | Apache-2.0 | |
| Hugging Face Vision Trainerhuggingface/skills | 11k | 1 repos | ~7.5k | Automated safety check: Pass | Apache-2.0 | |
| Huggingface LLM Trainerwaybarrios/opencode-power-pack | 533 | — | ~3k | Automated safety check: Pass | Apache-2.0 | |
| AI Newsletta-ai/skills | 149 | — | ~582 | Automated safety check: Pass | MIT |
huggingface/skills
Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.
huggingface/skills
Trains or fine-tunes language and vision models with TRL or Unsloth on Hugging Face Jobs cloud GPUs, then converts the results to GGUF.
huggingface/skills
Trains and fine-tunes object detection, image classification and SAM or SAM2 segmentation models on Hugging Face Jobs cloud GPUs and saves the results to the Hub.
waybarrios/opencode-power-pack
Train or fine-tune language models with TRL or Unsloth on Hugging Face Jobs, including SFT, DPO, GRPO, reward models, and GGUF conversion.
letta-ai/skills
Fetch and summarize recent AI news from curated RSS feeds (Hugging Face, VentureBeat, The Verge, OpenAI, Anthropic, DeepMind, etc.) and YouTube channels (Yannic Kilcher, Two Minute Papers, AI…
oracle/accelerated-data-science
Register, list, get, and manage LLM models in OCI AI Quick Actions (AQUA) using the ADS SDK.
awslabs/agent-plugins
Validates dataset formatting and quality for SageMaker model fine-tuning (SFT, DPO, or RLVR).
awslabs/agent-plugins
Evaluate, configure, and migrate workloads to AWS Lambda Managed Instances (LMI).
awslabs/agent-plugins
Selects a fine-tuning technique (SFT, DPO, RLVR, or RLAIF) for the user's use case and validates it against the selected model's available recipes.
awslabs/agent-plugins
Generate comprehensive issue reports from HyperPod clusters (EKS and Slurm) by collecting diagnostic logs and configurations for troubleshooting and AWS Support cases.
awslabs/agent-plugins
Diagnose performance issues on Amazon SageMaker HyperPod clusters — uneven NCCL bandwidth across nodes and poor filesystem throughput.
awslabs/agent-plugins
Diagnostic-only skill for Slurm scheduler and node-daemon issues on Amazon SageMaker HyperPod Slurm clusters.
Works with
Categories
Generates code that transforms datasets between ML schemas for model training or evaluation. Dataset Transformation is an agent skill from awslabs/agent-plugins, published by the product's own GitHub organization. Generates code that transforms datasets between ML schemas for model training or evaluation.
Dataset Transformation fits situations like: the user says transform; change the format.
Run `npx skills add awslabs/agent-plugins --skill dataset-transformation -a claude-code`. Or copy the skill folder (plugins/sagemaker-ai/skills/dataset-transformation in awslabs/agent-plugins) into .claude/skills/dataset-transformation in your project. Claude Code loads it when a task matches its description.
Run `npx skills add awslabs/agent-plugins --skill dataset-transformation -a codex`. Or copy the skill folder (plugins/sagemaker-ai/skills/dataset-transformation in awslabs/agent-plugins) into .agents/skills/dataset-transformation in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add awslabs/agent-plugins --skill dataset-transformation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/dataset-transformation, .gemini/skills/dataset-transformation, .github/skills/dataset-transformation and .opencode/skills/dataset-transformation in your project.
Going by SKILL.md and its folder, Dataset Transformation needs Python for the scripts in its folder and the command-line tools its instructions call (python3). Our summary lists: Python 3.
SKILL.md names 1 domain. As links in the text: docs.aws.amazon.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Dataset Transformation is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.5k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.8k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Dataset Transformation: SageMaker Serving Image Selection (huggingface/skills, 11k stars), Hugging Face LLM Trainer (huggingface/skills, 11k stars), Hugging Face Vision Trainer (huggingface/skills, 11k stars) and Huggingface LLM Trainer (waybarrios/opencode-power-pack, 533 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
awslabs (a GitHub organization, an official publisher) maintains it in awslabs/agent-plugins, which has 915 GitHub stars. The repository holds 33 skills in this directory. The repository was last updated on October 8, 2026.
Source: awslabs/agent-plugins on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.