Official agent skill

SageMaker Deployment Planner

by huggingface in huggingface/skills

Entry point for hosting a model on Amazon SageMaker: asks a few questions, picks a deployment pathway and hands off to the specialist skills.

OfficialApache-2.0Auto-check passedAI & LLM Engineering

Install SageMaker Deployment Planner

skills CLI
$ npx skills add huggingface/skills --skill hf-cloud-sagemaker-deployment-planner -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install huggingface/skills hf-cloud-sagemaker-deployment-planner --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/huggingface/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/hf-cloud-sagemaker-deployment-planner .claude/skills/hf-cloud-sagemaker-deployment-planner && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
hf-cloud-sagemaker-deployment-planner
GitHub stars
11k
Used in
1 other repo
Token cost
~2.1k tokens
SKILL.md length
1,056 words
Files
1
Skills in repo
25
Repo updated
First seen
Licence
Apache-2.0

At a glance

Entry point for hosting a model on Amazon SageMaker: asks a few questions, picks a deployment pathway and hands off to the specialist skills.

  • Works in 6 steps: Discovery — what is being deployed and… → Pathway selection — real-time /… → Context preflight —… → …
  • Deciding how to host a Hugging Face or fine-tuned model on AWS
  • SKILL.md covers Workflow phases, Discovery: ask only what you…, Pathway selection and Instance selection: check…, plus 1 more section
  • Calls aws

What it does

This planner owns the first two of six phases, discovery and pathway selection, and leaves the remaining phases to sibling skills for AWS context discovery, Python environment setup, IAM preflight, serving image selection and production defaults. Discovery asks only what is needed: which model (a Hugging Face ID, an S3 path or a name), how often it will be called, how much latency is tolerable and, when relevant, cost. The model type is usually inferred from the name, and the region comes from the context skill.

A table compares the pathways, which include real-time endpoints, real-time with scale to zero, serverless, async, batch and Bedrock CMI, with when each fits and when it does not. For instance, scale to zero suits sparse traffic, but a request after idle can take around nine minutes and every request during the wake fails with a 400. The skill covers text-generation LLMs, embedding models, rerankers, classifiers and diffusion models.

When your agent uses it

  • Deciding how to host a Hugging Face or fine-tuned model on AWS
  • Choosing between real-time and async inference for a new endpoint
  • Deploying an embedding model or reranker to SageMaker with sensible defaults
  • Starting a SageMaker deployment when you only know you want it running on AWS

Example prompts

  • “Deploy my fine-tuned model to SageMaker, I expect only a few requests an hour.”
  • “Host an embedding model from sentence-transformers on AWS for our search service.”
  • “I just want this running on AWS, so you pick the endpoint type.”
  • “Set up async inference on SageMaker for my text-to-image diffusion model.”

Requirements

  • An AWS account with SageMaker access
  • The sibling `hf-cloud-*` skills for the later phases

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Discovery — what is being deployed and what are the constraints (this skill)
  2. Pathway selection — real-time / serverless / async / batch / Bedrock CMI (this skill)
  3. Context preflight — hf-cloud-aws-context-discovery, then hf-cloud-python-env-setup
  4. IAM preflight — hf-cloud-sagemaker-iam-preflight
  5. Image selection — hf-cloud-serving-image-selection
  6. Deployment — hf-cloud-sagemaker-production-defaults

What it can do on your machine

Read from SKILL.md and the folder at commit ca0325b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • aws

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use aws, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

SageMaker Deployment Planner loads about 2.1k tokens when it runs. Until then it costs about 248 tokens; SKILL.md has 1,056 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~248
When it runs · the whole SKILL.md, loaded when a task matches
~2.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from huggingface/skills at commit ca0325b, republished under its Apache-2.0 licence (© huggingface). 1,056 words, ~2,146 tokens.

Download SKILL.mdSave it as .claude/skills/hf-cloud-sagemaker-deployment-planner/SKILL.md (or your agent's skills folder).
name
hf-cloud-sagemaker-deployment-planner
description
Plan and coordinate the deployment of a model to Amazon SageMaker AI. Use this skill whenever the user wants to deploy, host, serve, or expose a model on SageMaker or AWS — including phrases like "deploy a model", "host this LLM on AWS", "serve this embedding model", "deploy a reranker", "deploy a text-to-image / diffusion model", "host this for async inference", "create an endpoint", "serve my fine-tuned model", or any request that involves making a model available for inference on AWS. Use this even when the user is vague (e.g. "I just want to get this running on AWS, you figure it out"). Works for text-generation LLMs, embedding models, rerankers, classifiers, text-to-image / diffusion models — picks the right serving stack and chooses between real-time and async inference. This is the entry-point skill for SageMaker deployment work — it asks clarifying questions, picks a deployment pathway, and coordinates the other deployment skills.

SageMaker Deployment Planner

You are helping a user deploy a model to Amazon SageMaker. Most users invoking this skill want the model deployed with reasonable defaults, in as few questions as possible. Ask only what you need, recommend a pathway honestly, and hand off to the specialized skills.

Workflow phases

  1. Discovery — what is being deployed and what are the constraints (this skill)
  2. Pathway selection — real-time / serverless / async / batch / Bedrock CMI (this skill)
  3. Context preflight — hf-cloud-aws-context-discovery, then hf-cloud-python-env-setup
  4. IAM preflight — hf-cloud-sagemaker-iam-preflight
  5. Image selection — hf-cloud-serving-image-selection
  6. Deployment — hf-cloud-sagemaker-production-defaults

Phases 1–2 are this skill's job. The others activate when their patterns match.

Discovery: ask only what you need

You will eventually need to know:

  • What model: HuggingFace ID, S3 path to artifacts, or model name. If the user is vague ("the model I fine-tuned"), ask for the artifact location.
  • Model type: text-generation LLM, embedding/reranker, or other (classifier, NER, etc.). This determines the serving stack — usually inferable from the model name (anything ending in -embed-*, starting with BAAI/bge-, sentence-transformers/* etc. is embeddings; chat/instruct models are LLMs). Only ask if it's genuinely ambiguous.
  • Traffic shape: roughly how often will this be called?
  • Latency tolerance: interactive, near-real-time, or async?
  • Cost sensitivity: ask only if the user signals it or the traffic pattern is ambiguous.

Region comes from hf-cloud-aws-context-discovery — don't ask unless the user volunteers it.

Do not front-load all of these. A common minimal set is just: what model, and roughly how often will it be called? The model name usually settles the model-type question. That alone is often enough to narrow the pathway to two candidates. If the user already told you something, don't ask again.

Pathway selection

PathwayWhen it fitsWhen it does not
Real-time endpointSteady traffic, sub-second to few-second latency, always-onVery spiky or very sparse traffic (wastes money on idle)
Real-time, scale to zeroSparse or scheduled traffic, dev/test endpoints, and a client that tolerates a ~9 min first request after idleAny interactive SLA: every request during the wake fails with a 400
Serverless inferenceSpiky/intermittent, tolerates cold starts (~10s+), simpler modelsLLMs above a few B params (memory/cold-start limits), strict SLAs
Async inferenceLong inference (>60s), large payloads, queue-friendlyInteractive synchronous calls
Batch transformOffline scoring over a datasetAnything online or interactive
Bedrock Custom Model ImportWants Bedrock-compatible API, supported base family, weights onlyCustom inference logic, unsupported architectures

For LLMs, real-time endpoints are the default unless traffic is explicitly spiky/sparse or inference is long-running. Serverless looks attractive for "low traffic" cases but most LLMs exceed its memory limits.

For embeddings, real-time is again the default — but CPU instances are usually the right choice (much cheaper, fast enough for most embedding workloads). Don't reflexively recommend GPU instances for embedding models; ask hf-cloud-serving-image-selection to consider CPU variants if the model is small (<1B params) and traffic is moderate.

For text-to-image, video generation, or other long-inference workloads (>30s per request) where traffic is also bursty: async inference is the right answer. It supports genuine scale-to-zero between batches and queues requests via S3, so you don't pay for idle GPU. hf-cloud-sagemaker-production-defaults has a dedicated deploy_async.py for this.

Real-time, real-time scale-to-zero, and async are the three scripted pathways (deploy.py, deploy_ic.py, deploy_async.py in hf-cloud-sagemaker-production-defaults). Serverless, batch transform, and Bedrock Custom Model Import are not currently scripted — for those, hand the user off with a brief explanation rather than trying to deploy them through this workflow.

Scale to zero, real-time or async? Both reach zero and both make the first request after idle slow. Pick async when one inference can exceed the 60s InvokeEndpoint limit, when payloads are large, or when the client can accept an S3 result instead of a synchronous response. Pick real-time scale-to-zero when the client needs a normal synchronous HTTP response and can retry through the wake. Real-time scale-to-zero needs inference components; the plain real-time pathway cannot go below one instance.

If two pathways are both reasonable, say so in one sentence each and pick one. Don't bury the recommendation in options.

Show full SKILL.md (388 more words)Show less

Instance selection: check quota before recommending

Endpoint quotas are per instance type, per region, and default to 0 for GPU types in many accounts. Recommending an instance the account can't launch wastes a full deploy cycle on ResourceLimitExceeded. Check first:

bash
aws service-quotas list-service-quotas --service-code sagemaker --region <region> \
    --query "Quotas[?contains(QuotaName, 'for endpoint usage') && Value > \`0\`].[QuotaName, Value]" \
    --output table

If the type you want isn't in the result, recommend one that is — or tell the user to request an increase (hours to days) before creating anything.

If the call itself is denied, say so once and continue. The quota check is an optimization, not a gate: the deployment surfaces the real limit as ResourceLimitExceeded. Never stop the workflow, and never ask the user to change IAM, for a preflight check.

GPU family notes for the common 24 GB tier:

  • ml.g5.* (A10G) and ml.g6.* (L4) both work with current vLLM images when the gpu-3-1 AMI is set (see hf-cloud-serving-image-selection). g6 is the newer generation and slightly cheaper per hour; g5 has roughly double the memory bandwidth, which usually means better LLM token throughput. Pick whichever has quota; when both do, either is defensible — g5 for throughput, g6 for cost.
  • ml.g6e.* (L40S, 48 GB) when the model doesn't fit in 24 GB.

Once you have enough to recommend, state it plainly:

Based on what you've told me, I'd recommend a real-time endpoint on ml.g5.xlarge. The model is small enough that this is cost-effective, and your traffic pattern is steady enough that you won't be paying for idle. Alternative: serverless would be cheaper if traffic dries up for hours at a time, but Qwen3-0.6B is at the edge of serverless memory limits and cold starts would be 15–30s. Want me to proceed with the real-time endpoint?

Then wait for confirmation. The user should know what they're about to spend money on before you create anything.

The plan lives in the conversation — don't generate plan.yaml or similar artifacts unless explicitly asked.

Style

  • Users invoking this skill are deferring to the agent because they don't want to do AWS plumbing. Match that energy: efficient, not exhaustive.
  • One round of clarifying questions is usually enough. Three rounds is interrogation.
  • When you don't know something specific (current image URI, SDK API surface, quotas), check it rather than guess. Other skills handle the "how to check" details.
  • If the user pushes back on a recommendation, accept it. They know their constraints better than you do.

© huggingface, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/hf-cloud-sagemaker-deployment-planner of huggingface/skills.

Open the folder on GitHubat commit ca0325b

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in huggingface/skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

SageMaker Deployment Planner next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

SageMaker Deployment Planner compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
SageMaker Deployment Planner this skillhuggingface/skills11k1 repos~2.1kAutomated safety check: PassApache-2.0
Hf Cloud Serving Image Selectionwaybarrios/opencode-power-pack533—~4.3kAutomated safety check: PassApache-2.0
AWS AI MLaws/agent-toolkit-for-aws2.8k—~1.7kAutomated safety check: PassApache-2.0
Vllm Deploy K8svllm-project/vllm-skills103—~2kAutomated safety check: PassApache-2.0
Hf Cloud Sagemaker Production Defaultswaybarrios/opencode-power-pack533—~4.6kAutomated safety check: PassApache-2.0
Rtvi Byom PortingNVIDIA-AI-Blueprints/video-search-and-summarization1.9k—~1.4kAutomated safety check: PassApache-2.0

Similar skills

  • Hf Cloud Serving Image Selection

    waybarrios/opencode-power-pack

    Select and verify the current region-specific serving container URI for a SageMaker model deployment.

    533 GitHub stars~4.3k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • AWS AI ML

    aws/agent-toolkit-for-aws

    Official

    Selects, deploys, and customizes AI models on Amazon SageMaker.

    2.8k GitHub stars~1.7k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Vllm Deploy K8s

    vllm-project/vllm-skills

    Deploy vLLM to Kubernetes (K8s) with GPU support, health probes, and OpenAI-compatible API endpoint.

    103 GitHub stars~2k tokensUpdated 6 mo ago
    AI & LLM EngineeringAuto-check passed
  • Hf Cloud Sagemaker Production Defaults

    waybarrios/opencode-power-pack

    Implement a production SageMaker endpoint with autoscaling, CloudWatch alarms, and tags.

    533 GitHub stars~4.6k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Rtvi Byom Porting

    NVIDIA-AI-Blueprints/video-search-and-summarization

    A skill your agent uses when adding, debugging, or validating a bring-your-own VLM in VSS RT-VLM, including custom Hugging Face or NGC checkpoints, vLLM adapters or plugins, model shims, and…

    1.9k GitHub stars~1.4k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • LlamaGuard Content Moderation

    Orchestra-Research/AI-Research-SKILLs

    Uses Meta's LlamaGuard moderation model to screen prompts and model replies against six safety categories, with vLLM, FastAPI and NeMo Guardrails setups.

    13k GitHub starsUsed in 3 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed

More from huggingface/skills

All 25 skills in this repo
  • Official

    Finds or validates a usable SageMaker execution role before deploying or training, so scripts do not try to create IAM roles they lack permission to create.

    11k GitHub starsUsed in 1 repo~1.8k tokens
    Auto-check passed
  • Official

    Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.

    11k GitHub starsUsed in 1 repo~4.6k tokens
    Auto-check passed
  • Hugging Face LLM Trainer

    huggingface/skills

    Official

    Trains or fine-tunes language and vision models with TRL or Unsloth on Hugging Face Jobs cloud GPUs, then converts the results to GGUF.

    11k GitHub starsUsed in 3 repos~7.2k tokens
    Auto-check passed
  • Official

    Routes a sentence-transformers training task to the right model type and required reference docs and example scripts, covering bi-encoders, rerankers, sparse and multi-vector models.

    11k GitHub starsUsed in 1 repo~2.6k tokens
    Auto-check passed
  • Official

    Sets up an isolated Python environment with a supported interpreter and current boto3 before any SageMaker deployment, training or AWS automation code runs.

    11k GitHub starsUsed in 2 repos~1.7k tokens
    Auto-check passed
  • Official

    Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.

    11k GitHub starsUsed in 2 repos~1.6k tokens
    Auto-check passed

Questions about SageMaker Deployment Planner

What does SageMaker Deployment Planner do?

Entry point for hosting a model on Amazon SageMaker: asks a few questions, picks a deployment pathway and hands off to the specialist skills. This planner owns the first two of six phases, discovery and pathway selection, and leaves the remaining phases to sibling skills for AWS context discovery, Python environment setup, IAM preflight, serving image selection and production defaults. Discovery asks only what is needed: which model (a Hugging Face ID, an S3 path or a name), how often it will be called, how much latency is tolerable and, when relevant, cost.

When should I use SageMaker Deployment Planner?

SageMaker Deployment Planner fits situations like: deciding how to host a Hugging Face or fine-tuned model on AWS; choosing between real-time and async inference for a new endpoint; deploying an embedding model or reranker to SageMaker with sensible defaults; starting a SageMaker deployment when you only know you want it running on AWS.

How do I install SageMaker Deployment Planner in Claude Code?

Run `npx skills add huggingface/skills --skill hf-cloud-sagemaker-deployment-planner -a claude-code`. Or copy the skill folder (skills/hf-cloud-sagemaker-deployment-planner in huggingface/skills) into .claude/skills/hf-cloud-sagemaker-deployment-planner in your project. Claude Code loads it when a task matches its description.

How do I install SageMaker Deployment Planner in Codex?

Run `npx skills add huggingface/skills --skill hf-cloud-sagemaker-deployment-planner -a codex`. Or copy the skill folder (skills/hf-cloud-sagemaker-deployment-planner in huggingface/skills) into .agents/skills/hf-cloud-sagemaker-deployment-planner in your project. Codex loads it when a task matches its description.

Can I use SageMaker Deployment Planner in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add huggingface/skills --skill hf-cloud-sagemaker-deployment-planner -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/hf-cloud-sagemaker-deployment-planner, .gemini/skills/hf-cloud-sagemaker-deployment-planner, .github/skills/hf-cloud-sagemaker-deployment-planner and .opencode/skills/hf-cloud-sagemaker-deployment-planner in your project.

What does SageMaker Deployment Planner need to run?

Going by SKILL.md and its folder, SageMaker Deployment Planner needs the command-line tools its instructions call (aws). Our summary lists: An AWS account with SageMaker access; The sibling `hf-cloud-*` skills for the later phases.

Does SageMaker Deployment Planner access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is SageMaker Deployment Planner safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does SageMaker Deployment Planner use?

SageMaker Deployment Planner is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does SageMaker Deployment Planner use?

About 2.1k tokens (SKILL.md is roughly 8.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to SageMaker Deployment Planner?

Skills that share tags, products or a category with SageMaker Deployment Planner: Hf Cloud Serving Image Selection (waybarrios/opencode-power-pack, 533 stars), AWS AI ML (aws/agent-toolkit-for-aws, 2.8k stars), Vllm Deploy K8s (vllm-project/vllm-skills, 103 stars) and Hf Cloud Sagemaker Production Defaults (waybarrios/opencode-power-pack, 533 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains SageMaker Deployment Planner?

huggingface (a GitHub organization, an official publisher) maintains it in huggingface/skills, which has 11,142 GitHub stars. The repository holds 25 skills in this directory. The repository was last updated on October 1, 2026.

Source: huggingface/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.