Hf Cloud Serving Image Selection
waybarrios/opencode-power-pack
Select and verify the current region-specific serving container URI for a SageMaker model deployment.
Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.
$ npx skills add huggingface/skills --skill hf-cloud-serving-image-selection -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install huggingface/skills hf-cloud-serving-image-selection --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/huggingface/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/hf-cloud-serving-image-selection .claude/skills/hf-cloud-serving-image-selection && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "hf-cloud-serving-image-selection" agent skill from https://github.com/huggingface/skills/tree/main/skills/hf-cloud-serving-image-selection into .claude/skills/hf-cloud-serving-image-selection/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "hf-cloud-serving-image-selection", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/huggingface/skills/tree/main/skills/hf-cloud-serving-image-selectionType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add huggingface/skills --skill hf-cloud-serving-image-selection -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install huggingface/skills hf-cloud-serving-image-selection --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/huggingface/skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/hf-cloud-serving-image-selection .agents/skills/hf-cloud-serving-image-selection && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "hf-cloud-serving-image-selection" agent skill from https://github.com/huggingface/skills/tree/main/skills/hf-cloud-serving-image-selection into .agents/skills/hf-cloud-serving-image-selection/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "hf-cloud-serving-image-selection", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add huggingface/skills --skill hf-cloud-serving-image-selection -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install huggingface/skills hf-cloud-serving-image-selection --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/huggingface/skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/hf-cloud-serving-image-selection .cursor/skills/hf-cloud-serving-image-selection && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "hf-cloud-serving-image-selection" agent skill from https://github.com/huggingface/skills/tree/main/skills/hf-cloud-serving-image-selection into .cursor/skills/hf-cloud-serving-image-selection/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "hf-cloud-serving-image-selection", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/huggingface/skills.git --path skills/hf-cloud-serving-image-selection--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add huggingface/skills --skill hf-cloud-serving-image-selection -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install huggingface/skills hf-cloud-serving-image-selection --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/huggingface/skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/hf-cloud-serving-image-selection .gemini/skills/hf-cloud-serving-image-selection && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "hf-cloud-serving-image-selection" agent skill from https://github.com/huggingface/skills/tree/main/skills/hf-cloud-serving-image-selection into .gemini/skills/hf-cloud-serving-image-selection/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "hf-cloud-serving-image-selection", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install huggingface/skills hf-cloud-serving-image-selectionInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add huggingface/skills --skill hf-cloud-serving-image-selection -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/huggingface/skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/hf-cloud-serving-image-selection .github/skills/hf-cloud-serving-image-selection && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "hf-cloud-serving-image-selection" agent skill from https://github.com/huggingface/skills/tree/main/skills/hf-cloud-serving-image-selection into .github/skills/hf-cloud-serving-image-selection/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "hf-cloud-serving-image-selection", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add huggingface/skills --skill hf-cloud-serving-image-selection -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install huggingface/skills hf-cloud-serving-image-selection --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/huggingface/skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/hf-cloud-serving-image-selection .opencode/skills/hf-cloud-serving-image-selection && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "hf-cloud-serving-image-selection" agent skill from https://github.com/huggingface/skills/tree/main/skills/hf-cloud-serving-image-selection into .opencode/skills/hf-cloud-serving-image-selection/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "hf-cloud-serving-image-selection", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
hf-cloud-serving-image-selectionChooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.
A wrong container, stale tag or wrong AMI all surface as the same opaque health-check failure, so this skill makes the image choice explicit. When both a Hugging Face family (`huggingface-vllm`, `huggingface-vllm-omni`, `huggingface-sglang`, `tei`, `huggingface-pytorch-inference`) and a generic one (`vllm`, `vllm-omni`, `sglang`, `djl-inference`) can serve the model, the Hugging Face image is mandatory. A generic image is allowed only after a verified incompatibility, when no Hugging Face tag exists in the region, or when the image is on the known-broken list, and a newer version number alone does not count.
URIs are read from AWS's Deep Learning Containers catalog page, copying the example URL, substituting your region and passing it to `deploy.py --image-uri`. Some regions use different account IDs, so a region availability page is the fallback. A family missing from the catalog can be mirrored with `scripts/mirror_image.py`, and `references/model-to-image.md` maps models to images. The agent never hardcodes a URI from memory and never defaults to TGI.
3 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit ca0325b. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
curlpython3awsFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
huggingface.coapi.github.comraw.githubusercontent.comAlso links to:
aws.github.iogithub.comFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
HUGGING_FACE_HUB_TOKENHF_TOKENFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
SageMaker Serving Image Selection loads about 4.6k tokens when it runs, and up to ~7.3k if it reads all its reference files. Until then it costs about 253 tokens; SKILL.md has 2,135 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from huggingface/skills at commit ca0325b, republished under its Apache-2.0 licence (© huggingface). 2,135 words, ~4,609 tokens.
.claude/skills/hf-cloud-serving-image-selection/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.The serving container is the single thing most likely to break a SageMaker deployment that "looked correct on paper". Wrong container, stale tag, or the wrong AMI — all produce the same opaque Failed to pass health check error.
When both a HuggingFace-curated family (huggingface-vllm, huggingface-vllm-omni, huggingface-sglang, tei, huggingface-pytorch-inference) and a generic family (vllm, vllm-omni, sglang, djl-inference) can serve the model, the HuggingFace one is mandatory, not preferred. The only valid reasons to use a generic image:
A newer version number on the generic repo is not a reason. The AWS vllm repo often publishes a higher vLLM version than huggingface-vllm; an older-but-compatible huggingface-vllm tag still wins. "Latest vLLM" is not a requirement anyone stated — compatibility with the model is. If you fall back, record in the deployment log which of the three reasons applied.
Primary source: AWS's official Deep Learning Containers catalog.
URL: https://aws.github.io/deep-learning-containers/reference/available_images/
This page is AWS-maintained and lists every image family with example URIs, tags, CUDA versions, Python versions, and platform (SageMaker vs EC2/ECS/EKS). When picking a URI for a deployment, read it from this page directly — copy the example URL, substitute <region> with the user's region, and pass it to deploy.py --image-uri.
The example URLs use 763104351884 as the account ID for most regions. A few regions use different accounts (e.g. eu-south-1 uses 692866216735). Check the Region Availability page when in doubt.
Exception: none currently. Every image family used by this workflow is now on the AWS catalog page (TEI was added in late 2026). If you encounter a new family that isn't there, mirror it via mirror_image.py and pass the resulting URI directly.
| Model | Container family | How to get the URI |
|---|---|---|
| HuggingFace text-generation LLM (Llama, Qwen, Mistral, etc.) | HuggingFace vLLM | AWS catalog → "HuggingFace vLLM Inference" (ECR repo huggingface-vllm) |
| Same as above, multimodal | HuggingFace vLLM-Omni | AWS catalog → "HuggingFace vLLM-Omni Inference" (ECR repo huggingface-vllm-omni) |
| HuggingFace embeddings | TEI | AWS catalog → "HuggingFace Text Embeddings Inference" |
Encoder / cross-encoder rerankers (BERT-family *ForSequenceClassification) | TEI | Same as embeddings |
| Generative rerankers (causal-LM, e.g. Qwen3-Reranker) | HuggingFace vLLM | Same as text-generation LLMs — not TEI, see "Rerankers: TEI or vLLM?" |
| Text-to-image / diffusion (Stable Diffusion, FLUX) | DJL Inference | AWS catalog → "DJL Inference" — not HF Inference Toolkit, see "Known-broken images" |
| HuggingFace classifiers, NER, QA, summarization | HF Inference Toolkit (CPU) | AWS catalog → "HuggingFace PyTorch Inference"; GPU tags currently broken — see "Known-broken images" |
| User specifically wants SGLang | HuggingFace SGLang | AWS catalog → "HuggingFace SGLang Inference" |
No compatible huggingface-vllm tag (verified incompatibility or region gap — see "Rule zero") | vLLM (AWS) | AWS catalog → "vLLM" section — fallback only, never for version freshness |
| User specifically wants DJL-LMI | DJL Inference | AWS catalog → "DJL Inference" |
| Amazon Nova | SageMaker JumpStart | Use JumpStart, not raw endpoint creation |
| Custom inference code | BYOC | User provides URI |
HuggingFace-curated DLCs are mandatory when one is compatible (see "Rule zero"). huggingface-vllm is layered directly on the AWS vLLM DLC — identical SM_VLLM_* env contract and the same cu130 AMI rule — and adds current transformers, current huggingface_hub + hf_xet (avoids the XET-CDN 403 download failures older images hit), and HF performance defaults. It is also what SageMaker SDK v3 auto-routes to. The AWS vllm image is a compatibility escape hatch only; it usually shows a higher vLLM version than huggingface-vllm, and that is not a reason to pick it.
Do not use TGI. Text Generation Inference is archived. Models released after the archive (Qwen3 most famously) fail ping health checks on TGI. Use vLLM instead. (The SageMaker SDK v3 agrees: since PR #5960, June 2026, its ModelBuilder auto-routes text-generation to the HuggingFace vLLM DLC and multimodal tasks to HuggingFace vLLM-Omni.)
Full reasoning for each family in references/model-to-image.md.
"Reranker" covers two very different architectures, and picking wrong wastes a full endpoint-creation cycle (~20 min) before TEI rejects the model:
sentence-transformers rerankers) — BERT-family models with a classification head. config.json has architectures: [..ForSequenceClassification] on a TEI-supported encoder type. → TEI.config.json has architectures: [..ForCausalLM]. → HuggingFace vLLM, deployed exactly like a text-generation LLM. TEI will load the architecture then reject the classifier model type (Qwen3 support in TEI is embeddings-only). Invocation pattern (raw completions API, max_tokens=1, logprobs scoring) is in hf-cloud-sagemaker-production-defaults.Preflight before creating any resources — one HTTP GET settles it:
curl -s https://huggingface.co/<model-id>/raw/main/config.json
# "architectures": ["Qwen3ForCausalLM"] → vLLM
# "architectures": ["XLMRobertaForSequenceClassification"] → TEIFor TEI also confirm the (architecture, task) pair: an architecture appearing in TEI's supported list means embeddings support, not necessarily classification/reranking support.
Heads-up: SageMaker SDK v3 (PR #5960) routes the text-ranking task to TEI unconditionally — correct for cross-encoders, wrong for generative rerankers. Don't treat the SDK's routing as evidence that TEI can serve a given reranker.
For every family: read the URI from the AWS catalog page.
SageMaker for the platform column — newest within that family. Do not switch to another family's section because it lists a higher engine version (see "Rule zero")<region> with the user's region (from hf-cloud-aws-context-discovery)deploy.py --image-uri (real-time) or deploy_async.py --image-uri (async)The TEI catalog row lists two URIs — GPU (tei repo) and CPU (tei-cpu repo). Pick based on the instance type:
ml.g*, ml.p*, ml.inf* → GPU variantml.c*, ml.m*, ml.t* → CPU variantMixing them fails: CPU image on a GPU instance wastes hardware, GPU image on a CPU instance fails to start.
Note on the TEI account ID: the catalog page shows 683313688378 as the example account, but TEI is published from a different account namespace than the main AWS DLCs and the per-region account IDs vary. If 683313688378.dkr.ecr.<region>.amazonaws.com/tei:... returns an ECR pull error for a region other than us-east-1, check the Region Availability page for the correct account ID for that region.
vLLM DLC images with CUDA 13 or higher (current default: cu130) require setting InferenceAmiVersion=al2-ami-sagemaker-inference-gpu-3-1 on the ProductionVariant. This applies equally to huggingface-vllm and huggingface-vllm-omni (layered on the same cu130 base) and to the AWS vllm repo. Without it the container dies on startup with no CloudWatch logs ever created. The failure looks identical to many other things (account-level issues, quota, networking) and routinely sends people down wrong diagnostic paths.
Lookup table:
| Tag contains | InferenceAmiVersion to pass |
|---|---|
cu130 (or higher) | al2-ami-sagemaker-inference-gpu-3-1 |
cu129 or lower | (omit the flag; default AMI works) |
Rule of thumb: if the vLLM tag you picked contains cu130 or later, pass --inference-ami-version al2-ami-sagemaker-inference-gpu-3-1 to deploy.py. If a future CUDA version (cu140+) needs a different AMI, add a row to the table when AWS publishes the new image.
This is a vLLM-specific concern. TEI and HF Inference Toolkit images don't need an AMI override.
Both images share the same contract: configuration as environment variables on the SageMaker model definition, SM_VLLM_* mapped to vLLM CLI flags. The huggingface-vllm entrypoint additionally auto-detects the model when SM_VLLM_MODEL is unset — from /opt/ml/model if S3 artifacts are mounted, else from HF_MODEL_ID. For production, point SM_VLLM_MODEL at /opt/ml/model; loading directly from the Hub at runtime is the exception.
| Env var | Purpose | Notes |
|---|---|---|
SM_VLLM_MODEL | HF model ID (e.g. Qwen/Qwen3-0.6B) or /opt/ml/model if loading from S3 | — |
SM_VLLM_HOST | Must be 0.0.0.0 | Otherwise vLLM binds localhost only, ping fails, container dies before logs. Top cause of mystery failures with this image. |
SM_VLLM_TRUST_REMOTE_CODE | Whether model artifacts may execute custom Python code | Default false. Set true only for a specific architecture that cannot load without its repository code, after reviewing and pinning those artifacts. Enabling it allows model-supplied code to run inside the container. |
HUGGING_FACE_HUB_TOKEN | HF token | Required for gated models. |
Stage production model artifacts in an account-controlled S3 bucket, pass their URI as --model-s3-uri, and use SM_VLLM_MODEL=/opt/ml/model. This avoids downloading or executing repository code at runtime. Use a Hub model ID only when runtime Hub access is explicitly required; keep SM_VLLM_TRUST_REMOTE_CODE=false unless that architecture requires reviewed custom code.
| Env var | Purpose |
|---|---|
SM_VLLM_MAX_MODEL_LEN | Max sequence length — set this; defaults can be wrong for fine-tunes |
SM_VLLM_GPU_MEMORY_UTILIZATION | Float 0.0–1.0, ~0.9 reasonable |
SM_VLLM_TENSOR_PARALLEL_SIZE | GPU count for multi-GPU instances |
SM_VLLM_DTYPE | auto, bfloat16, float16 |
Any vLLM CLI flag works — uppercase, replace dashes with underscores, prepend SM_VLLM_.
Simpler env contract than vLLM:
| Env var | Purpose | Required |
|---|---|---|
HF_MODEL_ID | HF model ID (e.g. BAAI/bge-large-en-v1.5) or /opt/ml/model | Yes |
HF_TOKEN | HF auth token | Only for gated models |
MAX_BATCH_TOKENS | Max tokens per batch (default 16384) | No |
MAX_CLIENT_BATCH_SIZE | Max requests per client batch (default 32) | No |
No host-binding to configure, no trust-remote-code flag. The architectures TEI supports (BERT, CamemBERT, RoBERTa, XLM-RoBERTa, NomicBert, JinaBert, JinaCodeBert, Mistral, Qwen2/3, Gemma2/3, ModernBert) are baked into the image.
Critical and easy to get wrong:
| CUDA in image tag | Default AMI | With al2-ami-sagemaker-inference-gpu-3-1 |
|---|---|---|
| cu124 / cu128 | g5, g6, p5 all work | (not needed) |
| cu129 | g6, p5; g5 fails (driver mismatch → CannotStartContainerError) | expected to fix g5 (unverified) |
| cu130+ | fails everywhere — AMI flag is mandatory | g5, g6, p5 all work (cu130-on-g5 verified June 2026) |
The driver comes from the host AMI, not the instance family — so passing the gpu-3-1 AMI (which vLLM cu130 images require anyway) also makes ml.g5.* viable for cu129+ images.
SageMaker endpoints inside a VPC without a NAT gateway can't pull from public.ecr.aws. The deployment fails with an image-pull error that doesn't mention "VPC" or "egress".
For images on AWS's regional ECR (everything in the catalog): SageMaker reaches them through built-in routing, no NAT needed. Use the regional URI pattern (<account>.dkr.ecr.<region>.amazonaws.com/...), not the public.ecr.aws/... pattern.
For images requiring public.ecr.aws access (less common): mirror to a private ECR repo in your account with scripts/mirror_image.py (cross-platform; needs Docker + the aws CLI). Run it from the shell where the AWS CLI works.
# macOS / Linux
PRIVATE_URI=$(python3 scripts/mirror_image.py \
public.ecr.aws/deep-learning-containers/vllm:<tag> \
vllm-mirror)# Windows (PowerShell) — capture stdout into a variable
$PRIVATE_URI = python scripts\mirror_image.py `
public.ecr.aws/deep-learning-containers/vllm:<tag> vllm-mirrorThe page won't render / fetch returns junk: the catalog page is JavaScript-heavy and some fetch tools get an empty shell. Fallbacks, in order:
curl -s https://api.github.com/repos/aws/deep-learning-containers/contents/docs/src/data/huggingface-vllm
curl -s https://raw.githubusercontent.com/aws/deep-learning-containers/main/docs/src/data/huggingface-vllm/0.21.0-gpu-sagemaker.ymlhuggingface-vllm, huggingface-vllm-omni, huggingface-tei, vllm, djl-inference, ...).aws ecr describe-images --registry-id 763104351884 --repository-name huggingface-vllm \
--region <region> --query 'sort_by(imageDetails,&imagePushedAt)[-5:].imageTags' --output jsonA tag was just released and isn't on the page yet: rare; AWS updates the page on each release. Check the release notes above.
An architecture you need isn't supported by the listed image yet: for TEI specifically, you can mirror the upstream image from GHCR (ghcr.io/huggingface/text-embeddings-inference:<version>) into private ECR and pass the resulting URI directly to deploy.py --image-uri. Same mirror_image.py script.
| Image | Defect | Use instead |
|---|---|---|
huggingface-pytorch-inference GPU tags — all recent ones tested (PT 2.3–2.6, cu121/cu124, transformers 4.48–5.5.3) | ImportError: libtorch_cuda.so: undefined symbol: ncclCommResume at import torch. The NCCL bundled in the image is older than what torch links against — a packaging defect inside the container, on g5 and g6, regardless of AMI, model, or inference code. The MMS Java front-end keeps answering /ping, so the endpoint can reach InService while the Python worker crash-loops and serves nothing. | DJL Inference (bundles its own complete CUDA/NCCL stack) or BYOC. CPU tags are unaffected. |
Re-check when AWS publishes new huggingface-pytorch-inference GPU tags — remove the row once a fixed image is confirmed.
General fallback rule: when an HF DLC fails with CUDA/NCCL linker errors, switch to DJL Inference rather than iterating over sibling tags — the defect class is per-repo, not per-tag (three different tags were tried for the case above; all broken).
Related HF Hub gotcha: older DLCs can fail model download with 403 Forbidden from HF's XET CDN (their bundled huggingface_hub predates XET auth). Set HF_HUB_ENABLE_HF_TRANSFER=0 to force the standard download path, or pre-stage weights in S3.
Loading the model from HF Hub happens inside the container after the endpoint starts — expect 5–15+ minutes before InService even for small models, longer for multi-GB ones. A slow first boot is not a failure; don't tear down or re-diagnose before the deploy script's 30-minute wait expires.
For production or repeated deployments, pre-stage the weights in S3 and pass --model-s3-uri to deploy.py (the model then loads from /opt/ml/model) — faster, immune to Hub rate limits/outages, and no HUGGING_FACE_HUB_TOKEN needed at runtime.
© huggingface, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files (scripts, references) in skills/hf-cloud-serving-image-selection of huggingface/skills.
Open the folder on GitHubat commit ca0325b
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in huggingface/skills, which our catalogue first saw on October 7, 2026.
SageMaker Serving Image Selection next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| SageMaker Serving Image Selection this skillhuggingface/skills | 11k | 1 repos | ~4.6k | Automated safety check: Pass | Apache-2.0 | |
| Hf Cloud Serving Image Selectionwaybarrios/opencode-power-pack | 533 | — | ~4.3k | Automated safety check: Pass | Apache-2.0 | |
| Rtvi Byom PortingNVIDIA-AI-Blueprints/video-search-and-summarization | 1.9k | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | |
| Vllm Deploy K8svllm-project/vllm-skills | 103 | — | ~2k | Automated safety check: Pass | Apache-2.0 | |
| Hf Cloud Sagemaker Production Defaultswaybarrios/opencode-power-pack | 533 | — | ~4.6k | Automated safety check: Pass | Apache-2.0 | |
| AWS AI MLaws/agent-toolkit-for-aws | 2.8k | — | ~1.7k | Automated safety check: Pass | Apache-2.0 |
waybarrios/opencode-power-pack
Select and verify the current region-specific serving container URI for a SageMaker model deployment.
NVIDIA-AI-Blueprints/video-search-and-summarization
A skill your agent uses when adding, debugging, or validating a bring-your-own VLM in VSS RT-VLM, including custom Hugging Face or NGC checkpoints, vLLM adapters or plugins, model shims, and…
vllm-project/vllm-skills
Deploy vLLM to Kubernetes (K8s) with GPU support, health probes, and OpenAI-compatible API endpoint.
waybarrios/opencode-power-pack
Implement a production SageMaker endpoint with autoscaling, CloudWatch alarms, and tags.
aws/agent-toolkit-for-aws
Selects, deploys, and customizes AI models on Amazon SageMaker.
ericrisco/rsc-harness
A skill your agent uses when choosing an open-weight LLM and clearing it for use — which family and size fit the task, the hardware and the budget, and above all whether the license permits shipping.
huggingface/skills
Finds or validates a usable SageMaker execution role before deploying or training, so scripts do not try to create IAM roles they lack permission to create.
huggingface/skills
Trains or fine-tunes language and vision models with TRL or Unsloth on Hugging Face Jobs cloud GPUs, then converts the results to GGUF.
huggingface/skills
Routes a sentence-transformers training task to the right model type and required reference docs and example scripts, covering bi-encoders, rerankers, sparse and multi-vector models.
huggingface/skills
Sets up an isolated Python environment with a supported interpreter and current boto3 before any SageMaker deployment, training or AWS automation code runs.
huggingface/skills
Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.
huggingface/skills
Indexes research papers on the Hugging Face Hub from arXiv, links them to models and datasets, claims authorship and generates markdown research articles from templates.
Categories
Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones. A wrong container, stale tag or wrong AMI all surface as the same opaque health-check failure, so this skill makes the image choice explicit. When both a Hugging Face family (`huggingface-vllm`, `huggingface-vllm-omni`, `huggingface-sglang`, `tei`, `huggingface-pytorch-inference`) and a generic one (`vllm`, `vllm-omni`, `sglang`, `djl-inference`) can serve the model, the Hugging Face image is mandatory.
SageMaker Serving Image Selection fits situations like: deploying an LLM or fine-tuned model to a SageMaker endpoint; hosting an embedding model or reranker on SageMaker; replacing a hardcoded container URI in deployment code; debugging a SageMaker endpoint that fails its health check.
Run `npx skills add huggingface/skills --skill hf-cloud-serving-image-selection -a claude-code`. Or copy the skill folder (skills/hf-cloud-serving-image-selection in huggingface/skills) into .claude/skills/hf-cloud-serving-image-selection in your project. Claude Code loads it when a task matches its description.
Run `npx skills add huggingface/skills --skill hf-cloud-serving-image-selection -a codex`. Or copy the skill folder (skills/hf-cloud-serving-image-selection in huggingface/skills) into .agents/skills/hf-cloud-serving-image-selection in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add huggingface/skills --skill hf-cloud-serving-image-selection -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/hf-cloud-serving-image-selection, .gemini/skills/hf-cloud-serving-image-selection, .github/skills/hf-cloud-serving-image-selection and .opencode/skills/hf-cloud-serving-image-selection in your project.
Going by SKILL.md and its folder, SageMaker Serving Image Selection needs Python for the scripts in its folder, the command-line tools its instructions call (curl, python3 and aws) and credentials named HUGGING_FACE_HUB_TOKEN and HF_TOKEN. Our summary lists: An AWS account with SageMaker access in the target region; Network access to the AWS Deep Learning Containers catalog.
SKILL.md names 5 domains. In commands or code: huggingface.co, api.github.com and raw.githubusercontent.com; the agent is likely to contact these when it follows the instructions. As links in the text: aws.github.io and github.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
SageMaker Serving Image Selection is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.6k tokens (SKILL.md is roughly 18k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.7k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with SageMaker Serving Image Selection: Hf Cloud Serving Image Selection (waybarrios/opencode-power-pack, 533 stars), Rtvi Byom Porting (NVIDIA-AI-Blueprints/video-search-and-summarization, 1.9k stars), Vllm Deploy K8s (vllm-project/vllm-skills, 103 stars) and Hf Cloud Sagemaker Production Defaults (waybarrios/opencode-power-pack, 533 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
huggingface (a GitHub organization, an official publisher) maintains it in huggingface/skills, which has 11,142 GitHub stars. The repository holds 25 skills in this directory. The repository was last updated on October 1, 2026.
Source: huggingface/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.