SageMaker Serving Image Selection
huggingface/skills
Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.
Select and verify the current region-specific serving container URI for a SageMaker model deployment.
$ npx skills add waybarrios/opencode-power-pack --skill hf-cloud-serving-image-selection -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install waybarrios/opencode-power-pack hf-cloud-serving-image-selection --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/waybarrios/opencode-power-pack.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/hf-cloud-serving-image-selection .claude/skills/hf-cloud-serving-image-selection && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "hf-cloud-serving-image-selection" agent skill from https://github.com/waybarrios/opencode-power-pack/tree/main/skills/hf-cloud-serving-image-selection into .claude/skills/hf-cloud-serving-image-selection/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "hf-cloud-serving-image-selection", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/waybarrios/opencode-power-pack/tree/main/skills/hf-cloud-serving-image-selectionType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add waybarrios/opencode-power-pack --skill hf-cloud-serving-image-selection -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install waybarrios/opencode-power-pack hf-cloud-serving-image-selection --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/waybarrios/opencode-power-pack.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/hf-cloud-serving-image-selection .agents/skills/hf-cloud-serving-image-selection && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "hf-cloud-serving-image-selection" agent skill from https://github.com/waybarrios/opencode-power-pack/tree/main/skills/hf-cloud-serving-image-selection into .agents/skills/hf-cloud-serving-image-selection/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "hf-cloud-serving-image-selection", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add waybarrios/opencode-power-pack --skill hf-cloud-serving-image-selection -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install waybarrios/opencode-power-pack hf-cloud-serving-image-selection --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/waybarrios/opencode-power-pack.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/hf-cloud-serving-image-selection .cursor/skills/hf-cloud-serving-image-selection && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "hf-cloud-serving-image-selection" agent skill from https://github.com/waybarrios/opencode-power-pack/tree/main/skills/hf-cloud-serving-image-selection into .cursor/skills/hf-cloud-serving-image-selection/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "hf-cloud-serving-image-selection", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/waybarrios/opencode-power-pack.git --path skills/hf-cloud-serving-image-selection--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add waybarrios/opencode-power-pack --skill hf-cloud-serving-image-selection -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install waybarrios/opencode-power-pack hf-cloud-serving-image-selection --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/waybarrios/opencode-power-pack.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/hf-cloud-serving-image-selection .gemini/skills/hf-cloud-serving-image-selection && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "hf-cloud-serving-image-selection" agent skill from https://github.com/waybarrios/opencode-power-pack/tree/main/skills/hf-cloud-serving-image-selection into .gemini/skills/hf-cloud-serving-image-selection/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "hf-cloud-serving-image-selection", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install waybarrios/opencode-power-pack hf-cloud-serving-image-selectionInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add waybarrios/opencode-power-pack --skill hf-cloud-serving-image-selection -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/waybarrios/opencode-power-pack.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/hf-cloud-serving-image-selection .github/skills/hf-cloud-serving-image-selection && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "hf-cloud-serving-image-selection" agent skill from https://github.com/waybarrios/opencode-power-pack/tree/main/skills/hf-cloud-serving-image-selection into .github/skills/hf-cloud-serving-image-selection/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "hf-cloud-serving-image-selection", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add waybarrios/opencode-power-pack --skill hf-cloud-serving-image-selection -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install waybarrios/opencode-power-pack hf-cloud-serving-image-selection --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/waybarrios/opencode-power-pack.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/hf-cloud-serving-image-selection .opencode/skills/hf-cloud-serving-image-selection && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "hf-cloud-serving-image-selection" agent skill from https://github.com/waybarrios/opencode-power-pack/tree/main/skills/hf-cloud-serving-image-selection into .opencode/skills/hf-cloud-serving-image-selection/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "hf-cloud-serving-image-selection", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
hf-cloud-serving-image-selectionSelect and verify the current region-specific serving container URI for a SageMaker model deployment.
Hf Cloud Serving Image Selection is an agent skill from waybarrios/opencode-power-pack. Select and verify the current region-specific serving container URI for a SageMaker model deployment. Use after the deployment pathway is chosen and before endpoint code; never infer an image URI from memory.
Its SKILL.md is about 4.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including scripts and reference files (for example `references/model-to-image.md` and `scripts/mirror_image.py`).
It sits in AI & LLM Engineering, covering Model hubs and datasets, LLM inference and serving and Deployment. It works with Amazon SageMaker, Hugging Face, vLLM and Amazon Web Services. The repository describes itself as: 54 rigorous skills for Codex, OpenCode, and Pi: code review, security audit, feature development, frontend design, MCP tools, Hugging Face ML/training, and more. The licence is Apache-2.0.
3 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 9dccb6d. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
curlpython3awsFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
huggingface.coapi.github.comraw.githubusercontent.comAlso links to:
aws.github.iogithub.comFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
HUGGING_FACE_HUB_TOKENHF_TOKENFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Hf Cloud Serving Image Selection loads about 4.3k tokens when it runs, and up to ~7k if it reads all its reference files. Until then it costs about 60 tokens; SKILL.md has 2,059 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from waybarrios/opencode-power-pack at commit 9dccb6d, republished under its Apache-2.0 licence (© waybarrios). 2,059 words, ~4,286 tokens.
.claude/skills/hf-cloud-serving-image-selection/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.The serving container is the single thing most likely to break a SageMaker deployment that "looked correct on paper". Wrong container, stale tag, or the wrong AMI — all produce the same opaque Failed to pass health check error.
When both a HuggingFace-curated family (huggingface-vllm, huggingface-vllm-omni, huggingface-sglang, tei, huggingface-pytorch-inference) and a generic family (vllm, vllm-omni, sglang, djl-inference) can serve the model, the HuggingFace one is mandatory, not preferred. The only valid reasons to use a generic image:
A newer version number on the generic repo is not a reason. The AWS vllm repo often publishes a higher vLLM version than huggingface-vllm; an older-but-compatible huggingface-vllm tag still wins. "Latest vLLM" is not a requirement anyone stated — compatibility with the model is. If you fall back, record in the deployment log which of the three reasons applied.
Primary source: AWS's official Deep Learning Containers catalog.
URL: https://aws.github.io/deep-learning-containers/reference/available_images/
This page is AWS-maintained and lists every image family with example URIs, tags, CUDA versions, Python versions, and platform (SageMaker vs EC2/ECS/EKS). When picking a URI for a deployment, read it from this page directly — copy the example URL, substitute <region> with the user's region, and pass it to deploy.py --image-uri.
The example URLs use 763104351884 as the account ID for most regions. A few regions use different accounts (e.g. eu-south-1 uses 692866216735). Check the Region Availability page when in doubt.
Exception: none currently. Every image family used by this workflow is now on the AWS catalog page (TEI was added in late 2026). If you encounter a new family that isn't there, mirror it via mirror_image.py and pass the resulting URI directly.
| Model | Container family | How to get the URI |
|---|---|---|
| HuggingFace text-generation LLM (Llama, Qwen, Mistral, etc.) | HuggingFace vLLM | AWS catalog → "HuggingFace vLLM Inference" (ECR repo huggingface-vllm) |
| Same as above, multimodal | HuggingFace vLLM-Omni | AWS catalog → "HuggingFace vLLM-Omni Inference" (ECR repo huggingface-vllm-omni) |
| HuggingFace embeddings | TEI | AWS catalog → "HuggingFace Text Embeddings Inference" |
Encoder / cross-encoder rerankers (BERT-family *ForSequenceClassification) | TEI | Same as embeddings |
| Generative rerankers (causal-LM, e.g. Qwen3-Reranker) | HuggingFace vLLM | Same as text-generation LLMs — not TEI, see "Rerankers: TEI or vLLM?" |
| Text-to-image / diffusion (Stable Diffusion, FLUX) | DJL Inference | AWS catalog → "DJL Inference" — not HF Inference Toolkit, see "Known-broken images" |
| HuggingFace classifiers, NER, QA, summarization | HF Inference Toolkit (CPU) | AWS catalog → "HuggingFace PyTorch Inference"; GPU tags currently broken — see "Known-broken images" |
| User specifically wants SGLang | HuggingFace SGLang | AWS catalog → "HuggingFace SGLang Inference" |
No compatible huggingface-vllm tag (verified incompatibility or region gap — see "Rule zero") | vLLM (AWS) | AWS catalog → "vLLM" section — fallback only, never for version freshness |
| User specifically wants DJL-LMI | DJL Inference | AWS catalog → "DJL Inference" |
| Amazon Nova | SageMaker JumpStart | Use JumpStart, not raw endpoint creation |
| Custom inference code | BYOC | User provides URI |
HuggingFace-curated DLCs are mandatory when one is compatible (see "Rule zero"). huggingface-vllm is layered directly on the AWS vLLM DLC — identical SM_VLLM_* env contract and the same cu130 AMI rule — and adds current transformers, current huggingface_hub + hf_xet (avoids the XET-CDN 403 download failures older images hit), and HF performance defaults. It is also what SageMaker SDK v3 auto-routes to. The AWS vllm image is a compatibility escape hatch only; it usually shows a higher vLLM version than huggingface-vllm, and that is not a reason to pick it.
Do not use TGI. Text Generation Inference is archived. Models released after the archive (Qwen3 most famously) fail ping health checks on TGI. Use vLLM instead. (The SageMaker SDK v3 agrees: since PR #5960, June 2026, its ModelBuilder auto-routes text-generation to the HuggingFace vLLM DLC and multimodal tasks to HuggingFace vLLM-Omni.)
Full reasoning for each family in references/model-to-image.md.
"Reranker" covers two very different architectures, and picking wrong wastes a full endpoint-creation cycle (~20 min) before TEI rejects the model:
sentence-transformers rerankers) — BERT-family models with a classification head. config.json has architectures: [..ForSequenceClassification] on a TEI-supported encoder type. → TEI.config.json has architectures: [..ForCausalLM]. → HuggingFace vLLM, deployed exactly like a text-generation LLM. TEI will load the architecture then reject the classifier model type (Qwen3 support in TEI is embeddings-only). Invocation pattern (raw completions API, max_tokens=1, logprobs scoring) is in hf-cloud-sagemaker-production-defaults.Preflight before creating any resources — one HTTP GET settles it:
curl -s https://huggingface.co/<model-id>/raw/main/config.json
# "architectures": ["Qwen3ForCausalLM"] → vLLM
# "architectures": ["XLMRobertaForSequenceClassification"] → TEIFor TEI also confirm the (architecture, task) pair: an architecture appearing in TEI's supported list means embeddings support, not necessarily classification/reranking support.
Heads-up: SageMaker SDK v3 (PR #5960) routes the text-ranking task to TEI unconditionally — correct for cross-encoders, wrong for generative rerankers. Don't treat the SDK's routing as evidence that TEI can serve a given reranker.
For every family: read the URI from the AWS catalog page.
SageMaker for the platform column — newest within that family. Do not switch to another family's section because it lists a higher engine version (see "Rule zero")<region> with the user's region (from hf-cloud-aws-context-discovery)deploy.py --image-uri (real-time) or deploy_async.py --image-uri (async)The TEI catalog row lists two URIs — GPU (tei repo) and CPU (tei-cpu repo). Pick based on the instance type:
ml.g*, ml.p*, ml.inf* → GPU variantml.c*, ml.m*, ml.t* → CPU variantMixing them fails: CPU image on a GPU instance wastes hardware, GPU image on a CPU instance fails to start.
Note on the TEI account ID: the catalog page shows 683313688378 as the example account, but TEI is published from a different account namespace than the main AWS DLCs and the per-region account IDs vary. If 683313688378.dkr.ecr.<region>.amazonaws.com/tei:... returns an ECR pull error for a region other than us-east-1, check the Region Availability page for the correct account ID for that region.
vLLM DLC images with CUDA 13 or higher (current default: cu130) require setting InferenceAmiVersion=al2-ami-sagemaker-inference-gpu-3-1 on the ProductionVariant. This applies equally to huggingface-vllm and huggingface-vllm-omni (layered on the same cu130 base) and to the AWS vllm repo. Without it the container dies on startup with no CloudWatch logs ever created. The failure looks identical to many other things (account-level issues, quota, networking) and routinely sends people down wrong diagnostic paths.
Lookup table:
| Tag contains | InferenceAmiVersion to pass |
|---|---|
cu130 (or higher) | al2-ami-sagemaker-inference-gpu-3-1 |
cu129 or lower | (omit the flag; default AMI works) |
Rule of thumb: if the vLLM tag you picked contains cu130 or later, pass --inference-ami-version al2-ami-sagemaker-inference-gpu-3-1 to deploy.py. If a future CUDA version (cu140+) needs a different AMI, add a row to the table when AWS publishes the new image.
This is a vLLM-specific concern. TEI and HF Inference Toolkit images don't need an AMI override.
Both images share the same contract: configuration as environment variables on the SageMaker model definition, SM_VLLM_* mapped to vLLM CLI flags. The huggingface-vllm entrypoint additionally auto-detects the model when SM_VLLM_MODEL is unset — from /opt/ml/model if artifacts are mounted, else from HF_MODEL_ID — but setting SM_VLLM_MODEL explicitly works on both and is what our examples use.
| Env var | Purpose | Notes |
|---|---|---|
SM_VLLM_MODEL | HF model ID (e.g. Qwen/Qwen3-0.6B) or /opt/ml/model if loading from S3 | — |
SM_VLLM_HOST | Must be 0.0.0.0 | Otherwise vLLM binds localhost only, ping fails, container dies before logs. Top cause of mystery failures with this image. |
SM_VLLM_TRUST_REMOTE_CODE | true for Qwen and several recent architectures | Set unconditionally — downside negligible, upside is the model loads. |
HUGGING_FACE_HUB_TOKEN | HF token | Required for gated models. |
| Env var | Purpose |
|---|---|
SM_VLLM_MAX_MODEL_LEN | Max sequence length — set this; defaults can be wrong for fine-tunes |
SM_VLLM_GPU_MEMORY_UTILIZATION | Float 0.0–1.0, ~0.9 reasonable |
SM_VLLM_TENSOR_PARALLEL_SIZE | GPU count for multi-GPU instances |
SM_VLLM_DTYPE | auto, bfloat16, float16 |
Any vLLM CLI flag works — uppercase, replace dashes with underscores, prepend SM_VLLM_.
Simpler env contract than vLLM:
| Env var | Purpose | Required |
|---|---|---|
HF_MODEL_ID | HF model ID (e.g. BAAI/bge-large-en-v1.5) or /opt/ml/model | Yes |
HF_TOKEN | HF auth token | Only for gated models |
MAX_BATCH_TOKENS | Max tokens per batch (default 16384) | No |
MAX_CLIENT_BATCH_SIZE | Max requests per client batch (default 32) | No |
No host-binding to configure, no trust-remote-code flag. The architectures TEI supports (BERT, CamemBERT, RoBERTa, XLM-RoBERTa, NomicBert, JinaBert, JinaCodeBert, Mistral, Qwen2/3, Gemma2/3, ModernBert) are baked into the image.
Critical and easy to get wrong:
| CUDA in image tag | Default AMI | With al2-ami-sagemaker-inference-gpu-3-1 |
|---|---|---|
| cu124 / cu128 | g5, g6, p5 all work | (not needed) |
| cu129 | g6, p5; g5 fails (driver mismatch → CannotStartContainerError) | expected to fix g5 (unverified) |
| cu130+ | fails everywhere — AMI flag is mandatory | g5, g6, p5 all work (cu130-on-g5 verified June 2026) |
The driver comes from the host AMI, not the instance family — so passing the gpu-3-1 AMI (which vLLM cu130 images require anyway) also makes ml.g5.* viable for cu129+ images.
SageMaker endpoints inside a VPC without a NAT gateway can't pull from public.ecr.aws. The deployment fails with an image-pull error that doesn't mention "VPC" or "egress".
For images on AWS's regional ECR (everything in the catalog): SageMaker reaches them through built-in routing, no NAT needed. Use the regional URI pattern (<account>.dkr.ecr.<region>.amazonaws.com/...), not the public.ecr.aws/... pattern.
For images requiring public.ecr.aws access (less common): mirror to a private ECR repo in your account with scripts/mirror_image.py (cross-platform; needs Docker + the aws CLI). Run it from the shell where the AWS CLI works.
# macOS / Linux
PRIVATE_URI=$(python3 scripts/mirror_image.py \
public.ecr.aws/deep-learning-containers/vllm:<tag> \
vllm-mirror)# Windows (PowerShell) — capture stdout into a variable
$PRIVATE_URI = python scripts\mirror_image.py `
public.ecr.aws/deep-learning-containers/vllm:<tag> vllm-mirrorThe page won't render / fetch returns junk: the catalog page is JavaScript-heavy and some fetch tools get an empty shell. Fallbacks, in order:
curl -s https://api.github.com/repos/aws/deep-learning-containers/contents/docs/src/data/huggingface-vllm
curl -s https://raw.githubusercontent.com/aws/deep-learning-containers/main/docs/src/data/huggingface-vllm/0.21.0-gpu-sagemaker.ymlhuggingface-vllm, huggingface-vllm-omni, huggingface-tei, vllm, djl-inference, ...).aws ecr describe-images --registry-id 763104351884 --repository-name huggingface-vllm \
--region <region> --query 'sort_by(imageDetails,&imagePushedAt)[-5:].imageTags' --output jsonA tag was just released and isn't on the page yet: rare; AWS updates the page on each release. Check the release notes above.
An architecture you need isn't supported by the listed image yet: for TEI specifically, you can mirror the upstream image from GHCR (ghcr.io/huggingface/text-embeddings-inference:<version>) into private ECR and pass the resulting URI directly to deploy.py --image-uri. Same mirror_image.py script.
| Image | Defect | Use instead |
|---|---|---|
huggingface-pytorch-inference GPU tags — all recent ones tested (PT 2.3–2.6, cu121/cu124, transformers 4.48–5.5.3) | ImportError: libtorch_cuda.so: undefined symbol: ncclCommResume at import torch. The NCCL bundled in the image is older than what torch links against — a packaging defect inside the container, on g5 and g6, regardless of AMI, model, or inference code. The MMS Java front-end keeps answering /ping, so the endpoint can reach InService while the Python worker crash-loops and serves nothing. | DJL Inference (bundles its own complete CUDA/NCCL stack) or BYOC. CPU tags are unaffected. |
Re-check when AWS publishes new huggingface-pytorch-inference GPU tags — remove the row once a fixed image is confirmed.
General fallback rule: when an HF DLC fails with CUDA/NCCL linker errors, switch to DJL Inference rather than iterating over sibling tags — the defect class is per-repo, not per-tag (three different tags were tried for the case above; all broken).
Related HF Hub gotcha: older DLCs can fail model download with 403 Forbidden from HF's XET CDN (their bundled huggingface_hub predates XET auth). Set HF_HUB_ENABLE_HF_TRANSFER=0 to force the standard download path, or pre-stage weights in S3.
Loading the model from HF Hub happens inside the container after the endpoint starts — expect 5–15+ minutes before InService even for small models, longer for multi-GB ones. A slow first boot is not a failure; don't tear down or re-diagnose before the deploy script's 30-minute wait expires.
For production or repeated deployments, pre-stage the weights in S3 and pass --model-s3-uri to deploy.py (the model then loads from /opt/ml/model) — faster, immune to Hub rate limits/outages, and no HUGGING_FACE_HUB_TOKEN needed at runtime.
© waybarrios, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files (scripts, references) in skills/hf-cloud-serving-image-selection of waybarrios/opencode-power-pack.
Open the folder on GitHubat commit 9dccb6d
Hf Cloud Serving Image Selection next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Hf Cloud Serving Image Selection this skillwaybarrios/opencode-power-pack | 533 | — | ~4.3k | Automated safety check: Pass | Apache-2.0 | |
| SageMaker Serving Image Selectionhuggingface/skills | 11k | 1 repos | ~4.6k | Automated safety check: Pass | Apache-2.0 | |
| Rtvi Byom PortingNVIDIA-AI-Blueprints/video-search-and-summarization | 1.9k | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | |
| SageMaker Deployment Plannerhuggingface/skills | 11k | 1 repos | ~2.1k | Automated safety check: Pass | Apache-2.0 | |
| SageMaker Production Defaultshuggingface/skills | 11k | 1 repos | ~6.9k | Automated safety check: Pass | Apache-2.0 | |
| Python Environment Setup for SageMakerhuggingface/skills | 11k | 2 repos | ~1.7k | Automated safety check: Pass | Apache-2.0 |
huggingface/skills
Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.
NVIDIA-AI-Blueprints/video-search-and-summarization
A skill your agent uses when adding, debugging, or validating a bring-your-own VLM in VSS RT-VLM, including custom Hugging Face or NGC checkpoints, vLLM adapters or plugins, model shims, and…
huggingface/skills
Entry point for hosting a model on Amazon SageMaker: asks a few questions, picks a deployment pathway and hands off to the specialist skills.
huggingface/skills
Deploys SageMaker endpoints with autoscaling, CloudWatch alarms and tags on by default, using scripts for real-time, scale-to-zero and async setups.
huggingface/skills
Sets up an isolated Python environment with a supported interpreter and current boto3 before any SageMaker deployment, training or AWS automation code runs.
huggingface/skills
Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.
waybarrios/opencode-power-pack
Verify or select a SageMaker execution role before creating models, endpoints, or training jobs.
waybarrios/opencode-power-pack
Train or fine-tune language models with TRL or Unsloth on Hugging Face Jobs, including SFT, DPO, GRPO, reward models, and GGUF conversion.
waybarrios/opencode-power-pack
Train object-detection, image-classification, or SAM segmentation models on Hugging Face Jobs.
waybarrios/opencode-power-pack
Run CodeQL database creation and security queries, add data-extension models, or process CodeQL SARIF.
waybarrios/opencode-power-pack
Run Semgrep static analysis across a codebase, optionally using Semgrep Pro for cross-file taint analysis.
waybarrios/opencode-power-pack
Detects fail-open insecure defaults (hardcoded secrets, weak auth, permissive security) that allow apps to run insecurely in production.
Categories
Select and verify the current region-specific serving container URI for a SageMaker model deployment. Hf Cloud Serving Image Selection is an agent skill from waybarrios/opencode-power-pack. Select and verify the current region-specific serving container URI for a SageMaker model deployment.
Hf Cloud Serving Image Selection fits situations like: tasks that involve Model hubs and datasets; tasks that involve LLM inference and serving; tasks that involve Deployment.
Run `npx skills add waybarrios/opencode-power-pack --skill hf-cloud-serving-image-selection -a claude-code`. Or copy the skill folder (skills/hf-cloud-serving-image-selection in waybarrios/opencode-power-pack) into .claude/skills/hf-cloud-serving-image-selection in your project. Claude Code loads it when a task matches its description.
Run `npx skills add waybarrios/opencode-power-pack --skill hf-cloud-serving-image-selection -a codex`. Or copy the skill folder (skills/hf-cloud-serving-image-selection in waybarrios/opencode-power-pack) into .agents/skills/hf-cloud-serving-image-selection in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add waybarrios/opencode-power-pack --skill hf-cloud-serving-image-selection -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/hf-cloud-serving-image-selection, .gemini/skills/hf-cloud-serving-image-selection, .github/skills/hf-cloud-serving-image-selection and .opencode/skills/hf-cloud-serving-image-selection in your project.
Going by SKILL.md and its folder, Hf Cloud Serving Image Selection needs Python for the scripts in its folder, the command-line tools its instructions call (curl, python3 and aws) and credentials named HUGGING_FACE_HUB_TOKEN and HF_TOKEN. Our summary lists: Python 3; A credential in HUGGING_FACE_HUB_TOKEN.
SKILL.md names 5 domains. In commands or code: huggingface.co, api.github.com and raw.githubusercontent.com; the agent is likely to contact these when it follows the instructions. As links in the text: aws.github.io and github.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Hf Cloud Serving Image Selection is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.3k tokens (SKILL.md is roughly 17k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.7k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Hf Cloud Serving Image Selection: SageMaker Serving Image Selection (huggingface/skills, 11k stars), Rtvi Byom Porting (NVIDIA-AI-Blueprints/video-search-and-summarization, 1.9k stars), SageMaker Deployment Planner (huggingface/skills, 11k stars) and SageMaker Production Defaults (huggingface/skills, 11k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
waybarrios (a GitHub user) maintains it in waybarrios/opencode-power-pack, which has 533 GitHub stars. The repository holds 32 skills in this directory. The repository was last updated on October 6, 2026.
Source: waybarrios/opencode-power-pack on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.