Hf Cloud Serving Image Selection
waybarrios/opencode-power-pack
Select and verify the current region-specific serving container URI for a SageMaker model deployment.
Entry point for hosting a model on Amazon SageMaker: asks a few questions, picks a deployment pathway and hands off to the specialist skills.
$ npx skills add huggingface/skills --skill hf-cloud-sagemaker-deployment-planner -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install huggingface/skills hf-cloud-sagemaker-deployment-planner --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/huggingface/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/hf-cloud-sagemaker-deployment-planner .claude/skills/hf-cloud-sagemaker-deployment-planner && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "hf-cloud-sagemaker-deployment-planner" agent skill from https://github.com/huggingface/skills/tree/main/skills/hf-cloud-sagemaker-deployment-planner into .claude/skills/hf-cloud-sagemaker-deployment-planner/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "hf-cloud-sagemaker-deployment-planner", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/huggingface/skills/tree/main/skills/hf-cloud-sagemaker-deployment-plannerType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add huggingface/skills --skill hf-cloud-sagemaker-deployment-planner -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install huggingface/skills hf-cloud-sagemaker-deployment-planner --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/huggingface/skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/hf-cloud-sagemaker-deployment-planner .agents/skills/hf-cloud-sagemaker-deployment-planner && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "hf-cloud-sagemaker-deployment-planner" agent skill from https://github.com/huggingface/skills/tree/main/skills/hf-cloud-sagemaker-deployment-planner into .agents/skills/hf-cloud-sagemaker-deployment-planner/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "hf-cloud-sagemaker-deployment-planner", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add huggingface/skills --skill hf-cloud-sagemaker-deployment-planner -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install huggingface/skills hf-cloud-sagemaker-deployment-planner --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/huggingface/skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/hf-cloud-sagemaker-deployment-planner .cursor/skills/hf-cloud-sagemaker-deployment-planner && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "hf-cloud-sagemaker-deployment-planner" agent skill from https://github.com/huggingface/skills/tree/main/skills/hf-cloud-sagemaker-deployment-planner into .cursor/skills/hf-cloud-sagemaker-deployment-planner/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "hf-cloud-sagemaker-deployment-planner", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/huggingface/skills.git --path skills/hf-cloud-sagemaker-deployment-planner--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add huggingface/skills --skill hf-cloud-sagemaker-deployment-planner -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install huggingface/skills hf-cloud-sagemaker-deployment-planner --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/huggingface/skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/hf-cloud-sagemaker-deployment-planner .gemini/skills/hf-cloud-sagemaker-deployment-planner && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "hf-cloud-sagemaker-deployment-planner" agent skill from https://github.com/huggingface/skills/tree/main/skills/hf-cloud-sagemaker-deployment-planner into .gemini/skills/hf-cloud-sagemaker-deployment-planner/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "hf-cloud-sagemaker-deployment-planner", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install huggingface/skills hf-cloud-sagemaker-deployment-plannerInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add huggingface/skills --skill hf-cloud-sagemaker-deployment-planner -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/huggingface/skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/hf-cloud-sagemaker-deployment-planner .github/skills/hf-cloud-sagemaker-deployment-planner && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "hf-cloud-sagemaker-deployment-planner" agent skill from https://github.com/huggingface/skills/tree/main/skills/hf-cloud-sagemaker-deployment-planner into .github/skills/hf-cloud-sagemaker-deployment-planner/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "hf-cloud-sagemaker-deployment-planner", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add huggingface/skills --skill hf-cloud-sagemaker-deployment-planner -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install huggingface/skills hf-cloud-sagemaker-deployment-planner --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/huggingface/skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/hf-cloud-sagemaker-deployment-planner .opencode/skills/hf-cloud-sagemaker-deployment-planner && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "hf-cloud-sagemaker-deployment-planner" agent skill from https://github.com/huggingface/skills/tree/main/skills/hf-cloud-sagemaker-deployment-planner into .opencode/skills/hf-cloud-sagemaker-deployment-planner/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "hf-cloud-sagemaker-deployment-planner", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
hf-cloud-sagemaker-deployment-plannerEntry point for hosting a model on Amazon SageMaker: asks a few questions, picks a deployment pathway and hands off to the specialist skills.
This planner owns the first two of six phases, discovery and pathway selection, and leaves the remaining phases to sibling skills for AWS context discovery, Python environment setup, IAM preflight, serving image selection and production defaults. Discovery asks only what is needed: which model (a Hugging Face ID, an S3 path or a name), how often it will be called, how much latency is tolerable and, when relevant, cost. The model type is usually inferred from the name, and the region comes from the context skill.
A table compares the pathways, which include real-time endpoints, real-time with scale to zero, serverless, async, batch and Bedrock CMI, with when each fits and when it does not. For instance, scale to zero suits sparse traffic, but a request after idle can take around nine minutes and every request during the wake fails with a 400. The skill covers text-generation LLMs, embedding models, rerankers, classifiers and diffusion models.
6 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit ca0325b. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
awsFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use aws, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
SageMaker Deployment Planner loads about 2.1k tokens when it runs. Until then it costs about 248 tokens; SKILL.md has 1,056 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from huggingface/skills at commit ca0325b, republished under its Apache-2.0 licence (© huggingface). 1,056 words, ~2,146 tokens.
.claude/skills/hf-cloud-sagemaker-deployment-planner/SKILL.md (or your agent's skills folder).You are helping a user deploy a model to Amazon SageMaker. Most users invoking this skill want the model deployed with reasonable defaults, in as few questions as possible. Ask only what you need, recommend a pathway honestly, and hand off to the specialized skills.
hf-cloud-aws-context-discovery, then hf-cloud-python-env-setuphf-cloud-sagemaker-iam-preflighthf-cloud-serving-image-selectionhf-cloud-sagemaker-production-defaultsPhases 1–2 are this skill's job. The others activate when their patterns match.
You will eventually need to know:
-embed-*, starting with BAAI/bge-, sentence-transformers/* etc. is embeddings; chat/instruct models are LLMs). Only ask if it's genuinely ambiguous.Region comes from hf-cloud-aws-context-discovery — don't ask unless the user volunteers it.
Do not front-load all of these. A common minimal set is just: what model, and roughly how often will it be called? The model name usually settles the model-type question. That alone is often enough to narrow the pathway to two candidates. If the user already told you something, don't ask again.
| Pathway | When it fits | When it does not |
|---|---|---|
| Real-time endpoint | Steady traffic, sub-second to few-second latency, always-on | Very spiky or very sparse traffic (wastes money on idle) |
| Real-time, scale to zero | Sparse or scheduled traffic, dev/test endpoints, and a client that tolerates a ~9 min first request after idle | Any interactive SLA: every request during the wake fails with a 400 |
| Serverless inference | Spiky/intermittent, tolerates cold starts (~10s+), simpler models | LLMs above a few B params (memory/cold-start limits), strict SLAs |
| Async inference | Long inference (>60s), large payloads, queue-friendly | Interactive synchronous calls |
| Batch transform | Offline scoring over a dataset | Anything online or interactive |
| Bedrock Custom Model Import | Wants Bedrock-compatible API, supported base family, weights only | Custom inference logic, unsupported architectures |
For LLMs, real-time endpoints are the default unless traffic is explicitly spiky/sparse or inference is long-running. Serverless looks attractive for "low traffic" cases but most LLMs exceed its memory limits.
For embeddings, real-time is again the default — but CPU instances are usually the right choice (much cheaper, fast enough for most embedding workloads). Don't reflexively recommend GPU instances for embedding models; ask hf-cloud-serving-image-selection to consider CPU variants if the model is small (<1B params) and traffic is moderate.
For text-to-image, video generation, or other long-inference workloads (>30s per request) where traffic is also bursty: async inference is the right answer. It supports genuine scale-to-zero between batches and queues requests via S3, so you don't pay for idle GPU. hf-cloud-sagemaker-production-defaults has a dedicated deploy_async.py for this.
Real-time, real-time scale-to-zero, and async are the three scripted pathways (deploy.py, deploy_ic.py, deploy_async.py in hf-cloud-sagemaker-production-defaults). Serverless, batch transform, and Bedrock Custom Model Import are not currently scripted — for those, hand the user off with a brief explanation rather than trying to deploy them through this workflow.
Scale to zero, real-time or async? Both reach zero and both make the first request after idle slow. Pick async when one inference can exceed the 60s InvokeEndpoint limit, when payloads are large, or when the client can accept an S3 result instead of a synchronous response. Pick real-time scale-to-zero when the client needs a normal synchronous HTTP response and can retry through the wake. Real-time scale-to-zero needs inference components; the plain real-time pathway cannot go below one instance.
If two pathways are both reasonable, say so in one sentence each and pick one. Don't bury the recommendation in options.
Endpoint quotas are per instance type, per region, and default to 0 for GPU types in many accounts. Recommending an instance the account can't launch wastes a full deploy cycle on ResourceLimitExceeded. Check first:
aws service-quotas list-service-quotas --service-code sagemaker --region <region> \
--query "Quotas[?contains(QuotaName, 'for endpoint usage') && Value > \`0\`].[QuotaName, Value]" \
--output tableIf the type you want isn't in the result, recommend one that is — or tell the user to request an increase (hours to days) before creating anything.
If the call itself is denied, say so once and continue. The quota check is an optimization, not a gate: the deployment surfaces the real limit as ResourceLimitExceeded. Never stop the workflow, and never ask the user to change IAM, for a preflight check.
GPU family notes for the common 24 GB tier:
ml.g5.* (A10G) and ml.g6.* (L4) both work with current vLLM images when the gpu-3-1 AMI is set (see hf-cloud-serving-image-selection). g6 is the newer generation and slightly cheaper per hour; g5 has roughly double the memory bandwidth, which usually means better LLM token throughput. Pick whichever has quota; when both do, either is defensible — g5 for throughput, g6 for cost.ml.g6e.* (L40S, 48 GB) when the model doesn't fit in 24 GB.Once you have enough to recommend, state it plainly:
Based on what you've told me, I'd recommend a real-time endpoint on
ml.g5.xlarge. The model is small enough that this is cost-effective, and your traffic pattern is steady enough that you won't be paying for idle. Alternative: serverless would be cheaper if traffic dries up for hours at a time, but Qwen3-0.6B is at the edge of serverless memory limits and cold starts would be 15–30s. Want me to proceed with the real-time endpoint?
Then wait for confirmation. The user should know what they're about to spend money on before you create anything.
The plan lives in the conversation — don't generate plan.yaml or similar artifacts unless explicitly asked.
© huggingface, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/hf-cloud-sagemaker-deployment-planner of huggingface/skills.
Open the folder on GitHubat commit ca0325b
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in huggingface/skills, which our catalogue first saw on October 7, 2026.
SageMaker Deployment Planner next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| SageMaker Deployment Planner this skillhuggingface/skills | 11k | 1 repos | ~2.1k | Automated safety check: Pass | Apache-2.0 | |
| Hf Cloud Serving Image Selectionwaybarrios/opencode-power-pack | 533 | — | ~4.3k | Automated safety check: Pass | Apache-2.0 | |
| AWS AI MLaws/agent-toolkit-for-aws | 2.8k | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | |
| Vllm Deploy K8svllm-project/vllm-skills | 103 | — | ~2k | Automated safety check: Pass | Apache-2.0 | |
| Hf Cloud Sagemaker Production Defaultswaybarrios/opencode-power-pack | 533 | — | ~4.6k | Automated safety check: Pass | Apache-2.0 | |
| Rtvi Byom PortingNVIDIA-AI-Blueprints/video-search-and-summarization | 1.9k | — | ~1.4k | Automated safety check: Pass | Apache-2.0 |
waybarrios/opencode-power-pack
Select and verify the current region-specific serving container URI for a SageMaker model deployment.
aws/agent-toolkit-for-aws
Selects, deploys, and customizes AI models on Amazon SageMaker.
vllm-project/vllm-skills
Deploy vLLM to Kubernetes (K8s) with GPU support, health probes, and OpenAI-compatible API endpoint.
waybarrios/opencode-power-pack
Implement a production SageMaker endpoint with autoscaling, CloudWatch alarms, and tags.
NVIDIA-AI-Blueprints/video-search-and-summarization
A skill your agent uses when adding, debugging, or validating a bring-your-own VLM in VSS RT-VLM, including custom Hugging Face or NGC checkpoints, vLLM adapters or plugins, model shims, and…
Orchestra-Research/AI-Research-SKILLs
Uses Meta's LlamaGuard moderation model to screen prompts and model replies against six safety categories, with vLLM, FastAPI and NeMo Guardrails setups.
huggingface/skills
Finds or validates a usable SageMaker execution role before deploying or training, so scripts do not try to create IAM roles they lack permission to create.
huggingface/skills
Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.
huggingface/skills
Trains or fine-tunes language and vision models with TRL or Unsloth on Hugging Face Jobs cloud GPUs, then converts the results to GGUF.
huggingface/skills
Routes a sentence-transformers training task to the right model type and required reference docs and example scripts, covering bi-encoders, rerankers, sparse and multi-vector models.
huggingface/skills
Sets up an isolated Python environment with a supported interpreter and current boto3 before any SageMaker deployment, training or AWS automation code runs.
huggingface/skills
Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.
Categories
Entry point for hosting a model on Amazon SageMaker: asks a few questions, picks a deployment pathway and hands off to the specialist skills. This planner owns the first two of six phases, discovery and pathway selection, and leaves the remaining phases to sibling skills for AWS context discovery, Python environment setup, IAM preflight, serving image selection and production defaults. Discovery asks only what is needed: which model (a Hugging Face ID, an S3 path or a name), how often it will be called, how much latency is tolerable and, when relevant, cost.
SageMaker Deployment Planner fits situations like: deciding how to host a Hugging Face or fine-tuned model on AWS; choosing between real-time and async inference for a new endpoint; deploying an embedding model or reranker to SageMaker with sensible defaults; starting a SageMaker deployment when you only know you want it running on AWS.
Run `npx skills add huggingface/skills --skill hf-cloud-sagemaker-deployment-planner -a claude-code`. Or copy the skill folder (skills/hf-cloud-sagemaker-deployment-planner in huggingface/skills) into .claude/skills/hf-cloud-sagemaker-deployment-planner in your project. Claude Code loads it when a task matches its description.
Run `npx skills add huggingface/skills --skill hf-cloud-sagemaker-deployment-planner -a codex`. Or copy the skill folder (skills/hf-cloud-sagemaker-deployment-planner in huggingface/skills) into .agents/skills/hf-cloud-sagemaker-deployment-planner in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add huggingface/skills --skill hf-cloud-sagemaker-deployment-planner -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/hf-cloud-sagemaker-deployment-planner, .gemini/skills/hf-cloud-sagemaker-deployment-planner, .github/skills/hf-cloud-sagemaker-deployment-planner and .opencode/skills/hf-cloud-sagemaker-deployment-planner in your project.
Going by SKILL.md and its folder, SageMaker Deployment Planner needs the command-line tools its instructions call (aws). Our summary lists: An AWS account with SageMaker access; The sibling `hf-cloud-*` skills for the later phases.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
SageMaker Deployment Planner is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.1k tokens (SKILL.md is roughly 8.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with SageMaker Deployment Planner: Hf Cloud Serving Image Selection (waybarrios/opencode-power-pack, 533 stars), AWS AI ML (aws/agent-toolkit-for-aws, 2.8k stars), Vllm Deploy K8s (vllm-project/vllm-skills, 103 stars) and Hf Cloud Sagemaker Production Defaults (waybarrios/opencode-power-pack, 533 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
huggingface (a GitHub organization, an official publisher) maintains it in huggingface/skills, which has 11,142 GitHub stars. The repository holds 25 skills in this directory. The repository was last updated on October 1, 2026.
Source: huggingface/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.