Official agent skill

Model Evaluation

by awslabs in awslabs/agent-plugins

Generates python code that evaluates SageMaker models. An agent skill from awslabs/agent-plugins.

OfficialApache-2.0Auto-check passedAI & LLM Engineering

Install Model Evaluation

skills CLI
$ npx skills add awslabs/agent-plugins --skill model-evaluation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install awslabs/agent-plugins model-evaluation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/awslabs/agent-plugins.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/sagemaker-ai/skills/model-evaluation .claude/skills/model-evaluation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
model-evaluation
GitHub stars
915
Used in
1 other repo
Token cost
~1.3k tokens
SKILL.md length
706 words
Files
15 (incl. scripts, references)
Skills in repo
33
Repo updated
First seen
Licence
Apache-2.0

At a glance

Generates python code that evaluates SageMaker models. An agent skill from awslabs/agent-plugins.

  • Works in 2 steps: Determine evaluation type → Validate and hand off to evaluation…
  • The user says evaluate my model
  • SKILL.md covers Prerequisites, Principles, Scope and Evaluation Types, plus 1 more section
  • Runs Python scripts from its folder

What it does

Model Evaluation is an agent skill from awslabs/agent-plugins, published by the product's own GitHub organization. Generates python code that evaluates SageMaker models. Supports two evaluation types: LLM-as-Judge and Custom Scorer. Use when the user says "evaluate my model", "run a benchmark", "test model performance", "how did my model perform", "compare models", or other similar requests.

Its SKILL.md is about 1.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 17 other files, including scripts and reference files (for example `code_templates/custom_scorer_evaluator.py`, `code_templates/llmaaj_evaluator.py` and `references/code_output_guide.md`).

It sits in AI & LLM Engineering, covering Machine learning and LLM evaluation. It works with Amazon SageMaker, AWS Lambda and Python. The repository describes itself as: Agent Plugins for AWS equip AI coding agents with the skills to help you architect, deploy, and operate on AWS. The licence is Apache-2.0.

When your agent uses it

  • The user says evaluate my model
  • Run a benchmark
  • Test model performance
  • How did my model perform

Example prompts

  • “evaluate my model”
  • “run a benchmark”
  • “test model performance”
  • “/model-evaluation”

Requirements

  • Python 3

Workflow steps

2 steps, taken from the step headings in SKILL.md.

  1. Determine evaluation type
  2. Validate and hand off to evaluation workflow

What it can do on your machine

Read from SKILL.md and the folder at commit da51970. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 3 files in scripts/ (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Model Evaluation loads about 1.3k tokens when it runs, and up to ~15k if it reads all its reference files. Until then it costs about 74 tokens; SKILL.md has 706 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~74
When it runs · the whole SKILL.md, loaded when a task matches
~1.3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~15k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from awslabs/agent-plugins at commit da51970, republished under its Apache-2.0 licence (© awslabs). 706 words, ~1,325 tokens.

Download SKILL.mdSave it as .claude/skills/model-evaluation/SKILL.md (or your agent's skills folder). This skill also uses 14 other files; get the full folder from GitHub.
name
model-evaluation
description
Generates python code that evaluates SageMaker models. Supports two evaluation types: LLM-as-Judge and Custom Scorer. Use when the user says "evaluate my model", "run a benchmark", "test model performance", "how did my model perform", "compare models", or other similar requests.
metadata.version
3.0.0

Model Evaluation

Generate code that evaluates a SageMaker model.

Prerequisites

  • The SDK environment has been verified (SDK version, region, execution role). If not done, activate the sdk-getting-started skill first.

Principles

  1. One thing at a time. Each response advances exactly one decision. Never combine multiple questions in a single turn.
  2. Confirm before proceeding. Wait for the user to agree before moving to the next step.
  3. Don't read files until you need them. Only read reference files when you've reached the step that requires them.
  4. Don't ask what you already know. If the answer is in conversation history, workflow_state.json, plan.md, or any file you've already read — use it. Confirm if unsure, but don't re-ask.
  5. No narration. Share outcomes and ask questions. Keep responses short.
  6. No repetition. If you said something before a tool call, don't repeat it after.

Scope

This skill supports the evaluation feature for SageMaker Serverless Model Customization. It can evaluate any base or fine-tuned model supported by SageMaker serverless model customization — both OSS models (Llama, Mistral, Qwen, etc.) and Nova models.

Tell the user when the skill is activated:

"I can help evaluate any base or fine-tuned model supported by SageMaker serverless model customization."

If the user requests help evaluating a model that isn't supported by SageMaker serverless model customization, explain that it is not supported by this skill.

Evaluation Types

There are two evaluation types:

  • LLM-as-Judge — an LLM grades your model's responses. (OSS models only — not supported for Nova.)
  • Custom Scorer — programmatic evaluation via Lambda function (includes built-in math and code scorers). Works with both OSS and Nova models.

Workflow

Step 1: Determine evaluation type

Do you already know which evaluation type to use?

Check conversation history, plan.md, workflow_state.json, or anything else you've already read.

If yes: confirm with the user.

"It sounds like you want to run [evaluation type]. Is that right?"

⏸ Wait for confirmation. If confirmed → go to Step 2.

If no: ask.

"What kind of evaluation would you like to run? I support:

  1. LLM-as-Judge — an LLM grades your model's responses
  2. Custom Scorer — programmatic scoring (math, code, or your own logic)

Pick one, or say 'help me decide' if you're not sure."

⏸ Wait for user.

  • If user picks one → go to Step 2.
  • If user indicates uncertainty, by saying something like "help me decide," "whatever you think," "I'm not sure" → read references/evaluation-type-guide.md and follow its instructions. It will guide the user to a choice and then return here. You MUST NEVER make a recommendation to the user on eval type without reading references/evaluation-type-guide.md.
Show full SKILL.md (281 more words)Show less
Step 2: Validate and hand off to evaluation workflow

Before reading the reference file, validate that the chosen evaluation type is compatible with the user's situation. You may already know these answers from conversation context — don't ask if you don't need to.

LLM-as-Judge validation
  1. What model type are we evaluating? LLM-as-Judge is not supported for Nova models. To determine model type (if you don't already know it):
    • If you have the training job name or ARN, use the AWS MCP tool list-tags on the training job ARN and look for the sagemaker-studio:jumpstart-model-id tag. Contains "nova" → Nova. Anything else → OSS.
    • If you have a Model Package ARN, use the AWS MCP tool describe-model-package and check the model description or source tags.
    • If neither is available, ask the user.
  2. Does the user have an evaluation dataset? LLM-as-Judge requires one.
Custom Scorer validation
  1. Does the user have an evaluation dataset? Custom Scorer requires one. (Works with both OSS and Nova models, though for Nova only custom lambdas are supported.)

If validation fails, tell the user which requirement(s) aren't met and offer alternatives:

"[Evaluation type] won't work because [reason]."

If the failure reason was lack of an eval dataset, there's nothing we can do. Inform the user:

"Unfortunately all of the supported eval types require an eval dataset. I can't help you with model evaluation."

If the failure reason is something else, offer to help them pick a different evaluation type.

⏸ Wait for user.

If they say they do want help choosing a different eval type → read references/evaluation-type-guide.md.

If validation passes, read the corresponding reference file:

User choseRead
LLM-as-Judgereferences/llmaaj-evaluation.md
Custom Scorerreferences/custom-scorer-evaluation.md

Follow the reference file's instructions from the beginning.

© awslabs, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 14 other files (scripts, references) in plugins/sagemaker-ai/skills/model-evaluation of awslabs/agent-plugins.

  • SKILL.md
  • code_templates/custom_scorer_evaluator.py
  • code_templates/llmaaj_evaluator.py
  • references/code_output_guide.md
  • references/create-reward-function.md
  • references/custom-lambda-scorer.md
  • references/custom-scorer-evaluation.md
  • references/evaluation-type-guide.md
  • references/llmaaj-builtin-evaluation.md
  • references/llmaaj-custom-evaluation.md
  • references/llmaaj-evaluation.md
  • references/supported-judge-models.md
  • scripts/nova_reward_function_source_template.py
  • scripts/reward_function_source_template.py
  • scripts/validate_custom_metrics.py

Open the folder on GitHubat commit da51970

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in awslabs/agent-plugins, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Model Evaluation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Model Evaluation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Model Evaluation this skillawslabs/agent-plugins9151 repos~1.3kAutomated safety check: PassApache-2.0
LLM Benchmarking with lm-evaluation-harnessOrchestra-Research/AI-Research-SKILLs13k8 repos~3kAutomated safety check: PassMIT
Azure AI Projects Python SDKmicrosoft/skills3.1k6 repos~2.8kAutomated safety check: PassMIT
Python Environment Setup for SageMakerhuggingface/skills11k2 repos~1.7kAutomated safety check: PassApache-2.0
Failproof AI SDK IntegrationFailproofAI/failproofai5.3k—~6kAutomated safety check: PassCustom licence
Quality FlywheelGoogleCloudPlatform/vertex-ai-samples791—~2kAutomated safety check: PassApache-2.0

Similar skills

  • LLM Benchmarking with lm-evaluation-harness

    Orchestra-Research/AI-Research-SKILLs

    Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.

    13k GitHub starsUsed in 8 repos~3k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Reference for building on Microsoft Foundry with the azure-ai-projects Python SDK: project clients, versioned agents, evaluations, connections, datasets and indexes.

    3.1k GitHub starsUsed in 6 repos~2.8k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Sets up an isolated Python environment with a supported interpreter and current boto3 before any SageMaker deployment, training or AWS automation code runs.

    11k GitHub starsUsed in 2 repos~1.7k tokens
    AI & LLM EngineeringAuto-check passed
  • Failproof AI SDK Integration

    FailproofAI/failproofai

    Helps instrument a custom Python or TypeScript agent to record events for Failproof AI, verify what gets written, and run an evaluator worker that scores the runs.

    5.3k GitHub stars~6k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Quality Flywheel

    GoogleCloudPlatform/vertex-ai-samples

    Evaluate and improve GenAI models and agents using the Google GenAI Evaluation SDK.

    791 GitHub stars~2k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • PyTorch Lightning Training

    Orchestra-Research/AI-Research-SKILLs

    Shows how to organize PyTorch training with Lightning's LightningModule and Trainer, covering validation, DDP, callbacks and learning-rate scheduling.

    13k GitHub starsUsed in 7 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed

More from awslabs/agent-plugins

All 33 skills in this repo
  • Dataset Evaluation

    awslabs/agent-plugins

    Official

    Validates dataset formatting and quality for SageMaker model fine-tuning (SFT, DPO, or RLVR).

    915 GitHub starsUsed in 2 repos~1.3k tokens
    Auto-check passed
  • Dataset Transformation

    awslabs/agent-plugins

    Official

    Generates code that transforms datasets between ML schemas for model training or evaluation.

    915 GitHub starsUsed in 2 repos~3.5k tokens
    Auto-check passed
  • Finetuning Technique

    awslabs/agent-plugins

    Official

    Selects a fine-tuning technique (SFT, DPO, RLVR, or RLAIF) for the user's use case and validates it against the selected model's available recipes.

    915 GitHub starsUsed in 1 repo~604 tokens
    Auto-check passed
  • Hyperpod Issue Report

    awslabs/agent-plugins

    Official

    Generate comprehensive issue reports from HyperPod clusters (EKS and Slurm) by collecting diagnostic logs and configurations for troubleshooting and AWS Support cases.

    915 GitHub starsUsed in 1 repo~890 tokens
    Auto-check passed
  • Hyperpod Performance Debugger

    awslabs/agent-plugins

    Official

    Diagnose performance issues on Amazon SageMaker HyperPod clusters — uneven NCCL bandwidth across nodes and poor filesystem throughput.

    915 GitHub starsUsed in 1 repo~4.1k tokens
    Auto-check passed
  • Hyperpod Ssm

    awslabs/agent-plugins

    Official

    Remote command execution and file transfer on SageMaker HyperPod cluster nodes via AWS Systems Manager (SSM).

    915 GitHub starsUsed in 1 repo~1.3k tokens
    Auto-check: notes

Questions about Model Evaluation

What does Model Evaluation do?

Generates python code that evaluates SageMaker models. An agent skill from awslabs/agent-plugins. Model Evaluation is an agent skill from awslabs/agent-plugins, published by the product's own GitHub organization. Generates python code that evaluates SageMaker models.

When should I use Model Evaluation?

Model Evaluation fits situations like: the user says evaluate my model; run a benchmark; test model performance; how did my model perform.

How do I install Model Evaluation in Claude Code?

Run `npx skills add awslabs/agent-plugins --skill model-evaluation -a claude-code`. Or copy the skill folder (plugins/sagemaker-ai/skills/model-evaluation in awslabs/agent-plugins) into .claude/skills/model-evaluation in your project. Claude Code loads it when a task matches its description.

How do I install Model Evaluation in Codex?

Run `npx skills add awslabs/agent-plugins --skill model-evaluation -a codex`. Or copy the skill folder (plugins/sagemaker-ai/skills/model-evaluation in awslabs/agent-plugins) into .agents/skills/model-evaluation in your project. Codex loads it when a task matches its description.

Can I use Model Evaluation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add awslabs/agent-plugins --skill model-evaluation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/model-evaluation, .gemini/skills/model-evaluation, .github/skills/model-evaluation and .opencode/skills/model-evaluation in your project.

What does Model Evaluation need to run?

Going by SKILL.md and its folder, Model Evaluation needs Python for the scripts in its folder. Our summary lists: Python 3.

Does Model Evaluation access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Model Evaluation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Model Evaluation use?

Model Evaluation is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Model Evaluation use?

About 1.3k tokens (SKILL.md is roughly 5.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 13k tokens, read only when the agent opens those files.

What are the alternatives to Model Evaluation?

Skills that share tags, products or a category with Model Evaluation: LLM Benchmarking with lm-evaluation-harness (Orchestra-Research/AI-Research-SKILLs, 13k stars), Azure AI Projects Python SDK (microsoft/skills, 3.1k stars), Python Environment Setup for SageMaker (huggingface/skills, 11k stars) and Failproof AI SDK Integration (FailproofAI/failproofai, 5.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Model Evaluation?

awslabs (a GitHub organization, an official publisher) maintains it in awslabs/agent-plugins, which has 915 GitHub stars. The repository holds 33 skills in this directory. The repository was last updated on October 5, 2026.

Source: awslabs/agent-plugins on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.