Official agent skill

Finetuning

by awslabs in awslabs/agent-plugins

Generates code that fine-tunes a base model using SageMaker serverless training jobs.

OfficialApache-2.0Auto-check passedBackend & APIs

Install Finetuning

skills CLI
$ npx skills add awslabs/agent-plugins --skill finetuning -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install awslabs/agent-plugins finetuning --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/awslabs/agent-plugins.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/sagemaker-ai/skills/finetuning .claude/skills/finetuning && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
finetuning
GitHub stars
915
Used in
1 other repo
Token cost
~2.3k tokens
SKILL.md length
1,192 words
Files
14 (incl. scripts, references)
Skills in repo
33
Repo updated
First seen
Licence
Apache-2.0

At a glance

Generates code that fine-tunes a base model using SageMaker serverless training jobs.

  • Works in 6 steps: Code Generation Setup → RLVR Reward Function (for RLVR only,… → RLAIF (for RLAIF only, skip this section… → …
  • The user says start training
  • SKILL.md covers Code Generation Rules, User Communication Rules, 1. Code Generation Setup and 2. RLVR Reward Function (for…, plus 4 more sections
  • Runs Python scripts from its folder; calls python3 and python

What it does

Finetuning is an agent skill from awslabs/agent-plugins, published by the product's own GitHub organization. Generates code that fine-tunes a base model using SageMaker serverless training jobs. Use when the user says "start training", "fine-tune my model", "I'm ready to train", or when the plan reaches the finetuning step. Supports SFT, DPO, RLVR, and RLAIF trainers, including RLVR Lambda reward function and RLAIF custom prompt creation.

Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 17 other files, including scripts and reference files (for example `code_templates/dpo.py`, `code_templates/rlaif_builtin.py` and `code_templates/rlaif_custom_prompt.py`).

It sits in Backend & APIs, covering Fine-tuning and Serverless. It works with Amazon SageMaker. The repository describes itself as: Agent Plugins for AWS equip AI coding agents with the skills to help you architect, deploy, and operate on AWS. The licence is Apache-2.0.

When your agent uses it

  • The user says start training
  • Fine-tune my model
  • Im ready to train
  • The plan reaches the finetuning step

Example prompts

  • “start training”
  • “fine-tune my model”
  • “m ready to train”
  • “/finetuning”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Code Generation Setup
  2. RLVR Reward Function (for RLVR only, skip this section if technique is SFT or DPO)
  3. RLAIF (for RLAIF only, skip this section if technique is not RLAIF)
  4. EULA review and acceptance
  5. Post-Generation
  6. Continuous Customization

What it can do on your machine

Read from SKILL.md and the folder at commit da51970. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3
    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Finetuning loads about 2.3k tokens when it runs, and up to ~11k if it reads all its reference files. Until then it costs about 86 tokens; SKILL.md has 1,192 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~86
When it runs · the whole SKILL.md, loaded when a task matches
~2.3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~11k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from awslabs/agent-plugins at commit da51970, republished under its Apache-2.0 licence (© awslabs). 1,192 words, ~2,323 tokens.

Download SKILL.mdSave it as .claude/skills/finetuning/SKILL.md (or your agent's skills folder). This skill also uses 13 other files; get the full folder from GitHub.
name
finetuning
description
Generates code that fine-tunes a base model using SageMaker serverless training jobs. Use when the user says "start training", "fine-tune my model", "I'm ready to train", or when the plan reaches the finetuning step. Supports SFT, DPO, RLVR, and RLAIF trainers, including RLVR Lambda reward function and RLAIF custom prompt creation.
metadata.version
1.0.0

Prerequisites

Before starting this workflow, verify:

  1. A use_case_spec.md file exists

    • If missing: Activate the use-case-specification skill first, then resume
    • DON'T EVER offer to create a use case spec without activating the use-case-specification skill.
  2. A fine-tuning technique (SFT, DPO, RLVR, RLAIF, or CPT/RFT (for Nova)) and base model have already been selected

    • If missing: Activate the model-selection and/or finetuning-technique skills to collect what's missing, then resume
    • Don't make recommendations on the spot. You MUST activate the appropriate skill.
  3. A base model name available on SageMakerHub has been identified

    • If missing: Activate the model-selection skill to get it
    • Important: Only use the model name that model-selection retrieves, as it may differ from other commonly used names for the same model
  4. The SDK environment has been verified (SDK version, region, execution role)

    • If not done: Activate the sdk-getting-started skill first, then resume
  5. A training dataset uploaded to a bucket in the environment's default region.

    • If not met: Help the user upload the dataset to the correct S3

Critical Rules

Code Generation Rules

  • ✅ Use EXACTLY the imports shown in each code template
  • ❌ Do NOT add additional imports even if they seem helpful
  • ❌ Do NOT create variables before they're needed in that section
  • 📋 Copy the code structure precisely - no improvisation
  • 🎯 Follow the minimal code principle strictly
  • ✅ When writing code, make sure the indentation and f strings are correct

User Communication Rules

  • ❌ NEVER offer to move on to a downstream skill while training is in progress (logically impossible)
  • ❌ NEVER set ACCEPT_EULA to True without explicit user confirmation in the conversation
  • ✅ Always mention both the number AND title of sections you reference
  • ✅ If user asks how to run (notebook): If run_cell is available, offer to run it. Otherwise, tell them to run cells one by one (mention ipykernel requirement).
  • ✅ If user asks how to run (script): Tell them to run with python3 <script>.py

Workflow

1. Code Generation Setup

1.1 Directory Setup
  1. Identify project directory from conversation context
    • If unclear (multiple relevant directories exist) → Ask user which folder to use
    • If no project directory exists → activate the directory-management skill to set one up

⏸ Wait for user.

1.2 Select Code Template

Read references/code_output_guide.md for output format rules, then read the code template matching the finetuning strategy:

  • SFT → code_templates/sft.py
  • DPO → code_templates/dpo.py
  • RLVR → code_templates/rlvr.py
  • RLAIF with built-in rewards → code_templates/rlaif_builtin.py
  • RLAIF with custom prompt → code_templates/rlaif_custom_prompt.py

The template is a Python file where each # Cell N: Label comment marks the start of a new section. Split on these markers — everything between one marker and the next becomes one unit of output.

1.3 Generate Code
  1. Write the code from the template following the rules in code_output_guide.md
  2. Use same order, dependencies, and imports as the template
  3. DO NOT improvise or add extra code
  4. If the model is NOT a Meta/Llama model (model ID does NOT start with meta-):
    • Omit the ACCEPT_EULA = False line from the config cell
    • Omit the accept_eula=ACCEPT_EULA, line from the trainer call
  5. If the model is from the Nova family, omit any code containing max_epochs or lr_warmup_steps_ratio from the Configure Trainer section and the Hyperparameter Overrides section
1.4 Auto-Generate Configuration Values

In the 'Setup & Credentials' cell, populate:

  1. BASE_MODEL

    • Use the exact SageMakerHub model name from context
  2. MODEL_PACKAGE_GROUP_NAME

    • Generate from use case (read use_case_spec.md if needed)
    • Format rules:
      • Lowercase, alphanumeric with hyphens only
      • 1-63 characters
      • Pattern: [a-zA-Z0-9](-*[a-zA-Z0-9]){0,62}
      • Example: "Customer Support Chatbot" → customer-support-chatbot-v1
  3. Save notebook

2. RLVR Reward Function (for RLVR only, skip this section if technique is SFT or DPO)

2.1 Check Reward Function Status
  • Ask if user has a reward function already, or would like help creating one.
    • If user says they have one → Ask for the SageMaker Hub Evaluator ARN. Only proceed to Section 2.3 once the user provides a valid Evaluator ARN. If they don't have it registered as a SageMaker Hub Evaluator, continue to 2.2.
    • If user says they do not have one → Continue to 2.2
2.2 Generate Reward Function From Template
  1. Follow workflow in references/rlvr_reward_function.md section "Helping Users Create Custom Reward Functions"
2.3 Set CUSTOM_REWARD_FUNCTION value
  1. Set the value for CUSTOM_REWARD_FUNCTION in the Notebook with the ARN of the reward function (either given directly by the user, or from the function generation code as evaluator.arn).

3. RLAIF (for RLAIF only, skip this section if technique is not RLAIF)

Read references/rlaif_guide.md and follow its instructions.

Show full SKILL.md (470 more words)Show less

4. EULA review and acceptance

  1. Look up the official license link for the selected base model from references/eula_links.md
  2. Display the license to the user following the phrasing in references/eula_links.md. For OSS models: "This model is licensed under {License}. Please review the license terms here: {URL}." For Nova models: "This model is subject to the AWS Service Terms: {URL}."
  3. Check if the selected base model is a Meta/Llama model (model ID starts with meta-)
    • If Meta/Llama: Tell the user they must read and agree to the EULA before using this model. Ask: "Do you accept the license terms? (yes/no)". If the user confirms, set ACCEPT_EULA = True and uncomment accept_eula=ACCEPT_EULA in the generated notebook. If the user declines, leave ACCEPT_EULA = False and warn that training will fail without acceptance.
    • If non-Meta: Inform the user of the license for their awareness. No code-level action needed — the ACCEPT_EULA variable and accept_eula parameter should already be omitted from the notebook (see Step 1.3).

5. Post-Generation

After generating the code, offer to run it. Training can take hours depending on your dataset and model.

Notebook mode: If run_cell is available, offer to run the cells. Otherwise tell the user to run cells themselves.

Script mode: Present the user with options:

"Would you like me to:

  1. Leave it to you — run with python scripts/[script_name]
  2. Run it and wait until it's done
  3. Start it but don't wait — we can check status later"
  • Option 1: Done. Wait for user to come back.
  • Option 2: Execute the script as-is. trainer.train(wait=True) blocks until complete. Report final status.
  • Option 3: Change wait=True to wait=False in the script, execute, report the training job name.

Checking status:

  • describe-training-job --training-job-name NAME → TrainingJobStatus, FailureReason, SecondaryStatusTransitions
  • For model package ARN after completion: list-model-packages --model-package-group-name GROUP_NAME --sort-by CreationTime --sort-order Descending --max-results 1

Showing results after completion:

  • Use scripts/mlflow_reference.py as the pattern to query MLflow metrics
  • Present loss by epoch as a text table (total_loss, val_eval_total_loss for SFT; rewards/margins for DPO; critic/rewards/mean for RLVR)

CRITICAL:

  • DON'T suggest moving to next steps before training completes
  • DON'T elaborate on the next steps unless the user specifically asks you about them.

6. Continuous Customization

If the user wants to finetune a model they had already customized, follow the instructions in references/continuous_customization.md


References

  • rlvr_reward_function.md - Lambda reward function creation guide (RLVR only)
  • templates/rlvr_reward_function_source_template.py - Lambda reward function source template for open-weights models (RLVR only)
  • templates/nova_rlvr_reward_function_source_template.py - Lambda reward function source template for Nova 2.0 Lite (RLVR only)
  • code_templates/sft.py - Complete notebook template for Supervised Fine-Tuning (OSS path)
  • code_templates/dpo.py - Complete notebook template for Direct Preference Optimization (OSS path)
  • code_templates/rlvr.py - Complete notebook template for Reinforcement Learning from Verifiable Rewards (OSS path)
  • references/continuous_customization.md - Instructions on fine-tuning an already fine-tuned model.
  • rlaif_guide.md - instructions on RLAIF finetuning options
  • rlaif_builtin.py - Code template for RLAIF with built-in judge prompt
  • rlaif_custom_prompt.py - Code template for RLAIF with custom judge prompt

© awslabs, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 13 other files (scripts, references) in plugins/sagemaker-ai/skills/finetuning of awslabs/agent-plugins.

  • SKILL.md
  • code_templates/dpo.py
  • code_templates/rlaif_builtin.py
  • code_templates/rlaif_custom_prompt.py
  • code_templates/rlvr.py
  • code_templates/sft.py
  • references/code_output_guide.md
  • references/continuous_customization.md
  • references/eula_links.md
  • references/rlaif_guide.md
  • references/rlvr_reward_function.md
  • scripts/mlflow_reference.py
  • templates/nova_rlvr_reward_function_source_template.py
  • templates/rlvr_reward_function_source_template.py

Open the folder on GitHubat commit da51970

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in awslabs/agent-plugins, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Finetuning next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Finetuning compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Finetuning this skillawslabs/agent-plugins9151 repos~2.3kAutomated safety check: PassApache-2.0
AWS AI MLaws/agent-toolkit-for-aws2.8k—~1.7kAutomated safety check: PassApache-2.0
ModalK-Dense-AI/scientific-agent-skills48k1 repos~4.5kAutomated safety check: NotesApache-2.0
Serverless ModalAI4Scientist/nano-scientist1284 repos~3.1kAutomated safety check: NotesNone
Runpodericrisco/rsc-harness167—~2.8kAutomated safety check: PassMIT
AWS Harnesshoodini/ai-agents-skills281—~4.2kAutomated safety check: NotesNone

Similar skills

  • AWS AI ML

    aws/agent-toolkit-for-aws

    Official

    Selects, deploys, and customizes AI models on Amazon SageMaker.

    2.8k GitHub stars~1.7k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Modal

    K-Dense-AI/scientific-agent-skills

    Modal is a serverless cloud platform for running Python on demand, including on-demand GPUs.

    48k GitHub starsUsed in 1 repo~4.5k tokens
    Backend & APIsAuto-check: notes
  • Serverless Modal

    AI4Scientist/nano-scientist

    Run GPU workloads on Modal — training, fine-tuning, inference, batch processing.

    128 GitHub starsUsed in 4 repos~3.1k tokens
    Backend & APIsAuto-check: notes
  • Runpod

    ericrisco/rsc-harness

    A skill your agent uses when running GPU compute on RunPod and deciding between Pods (hourly, always-on) and Serverless (per-second, autoscaling) for training, fine-tuning or inference — serverless…

    167 GitHub stars~2.8k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • AWS Harness

    hoodini/ai-agents-skills

    Build a new AI agent on AWS and deploy it easily, OR wrap and deploy an agent you already have, using the Amazon Bedrock AgentCore harness.

    281 GitHub stars~4.2k tokensUpdated 2 mo ago
    Backend & APIsAuto-check: notes
  • Together Reference Architecture

    jeremylongshore/tons-of-skills-marketplace

    Design a production Together AI service with a typed provider boundary, policy-based model routing, serverless and dedicated lanes, batch workers, telemetry, budgets, and reversible degradation.

    2.8k GitHub stars~1.1k tokensUpdated today
    Backend & APIsAuto-check passed

More from awslabs/agent-plugins

All 33 skills in this repo
  • Dataset Evaluation

    awslabs/agent-plugins

    Official

    Validates dataset formatting and quality for SageMaker model fine-tuning (SFT, DPO, or RLVR).

    915 GitHub starsUsed in 2 repos~1.3k tokens
    Auto-check passed
  • Dataset Transformation

    awslabs/agent-plugins

    Official

    Generates code that transforms datasets between ML schemas for model training or evaluation.

    915 GitHub starsUsed in 2 repos~3.5k tokens
    Auto-check passed
  • Finetuning Technique

    awslabs/agent-plugins

    Official

    Selects a fine-tuning technique (SFT, DPO, RLVR, or RLAIF) for the user's use case and validates it against the selected model's available recipes.

    915 GitHub starsUsed in 1 repo~604 tokens
    Auto-check passed
  • Hyperpod Issue Report

    awslabs/agent-plugins

    Official

    Generate comprehensive issue reports from HyperPod clusters (EKS and Slurm) by collecting diagnostic logs and configurations for troubleshooting and AWS Support cases.

    915 GitHub starsUsed in 1 repo~890 tokens
    Auto-check passed
  • Hyperpod Performance Debugger

    awslabs/agent-plugins

    Official

    Diagnose performance issues on Amazon SageMaker HyperPod clusters — uneven NCCL bandwidth across nodes and poor filesystem throughput.

    915 GitHub starsUsed in 1 repo~4.1k tokens
    Auto-check passed
  • Hyperpod Ssm

    awslabs/agent-plugins

    Official

    Remote command execution and file transfer on SageMaker HyperPod cluster nodes via AWS Systems Manager (SSM).

    915 GitHub starsUsed in 1 repo~1.3k tokens
    Auto-check: notes

Questions about Finetuning

What does Finetuning do?

Generates code that fine-tunes a base model using SageMaker serverless training jobs. Finetuning is an agent skill from awslabs/agent-plugins, published by the product's own GitHub organization. Generates code that fine-tunes a base model using SageMaker serverless training jobs.

When should I use Finetuning?

Finetuning fits situations like: the user says start training; fine-tune my model; im ready to train; the plan reaches the finetuning step.

How do I install Finetuning in Claude Code?

Run `npx skills add awslabs/agent-plugins --skill finetuning -a claude-code`. Or copy the skill folder (plugins/sagemaker-ai/skills/finetuning in awslabs/agent-plugins) into .claude/skills/finetuning in your project. Claude Code loads it when a task matches its description.

How do I install Finetuning in Codex?

Run `npx skills add awslabs/agent-plugins --skill finetuning -a codex`. Or copy the skill folder (plugins/sagemaker-ai/skills/finetuning in awslabs/agent-plugins) into .agents/skills/finetuning in your project. Codex loads it when a task matches its description.

Can I use Finetuning in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add awslabs/agent-plugins --skill finetuning -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/finetuning, .gemini/skills/finetuning, .github/skills/finetuning and .opencode/skills/finetuning in your project.

What does Finetuning need to run?

Going by SKILL.md and its folder, Finetuning needs Python for the scripts in its folder and the command-line tools its instructions call (python3 and python). Our summary lists: Python 3.

Does Finetuning access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Finetuning safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Finetuning use?

Finetuning is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Finetuning use?

About 2.3k tokens (SKILL.md is roughly 9.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 8.5k tokens, read only when the agent opens those files.

What are the alternatives to Finetuning?

Skills that share tags, products or a category with Finetuning: AWS AI ML (aws/agent-toolkit-for-aws, 2.8k stars), Modal (K-Dense-AI/scientific-agent-skills, 48k stars), Serverless Modal (AI4Scientist/nano-scientist, 128 stars) and Runpod (ericrisco/rsc-harness, 167 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Finetuning?

awslabs (a GitHub organization, an official publisher) maintains it in awslabs/agent-plugins, which has 915 GitHub stars. The repository holds 33 skills in this directory. The repository was last updated on October 5, 2026.

Source: awslabs/agent-plugins on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.