SageMaker Serving Image Selection
huggingface/skills
Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.
Implement a production SageMaker endpoint with autoscaling, CloudWatch alarms, and tags.
$ npx skills add waybarrios/opencode-power-pack --skill hf-cloud-sagemaker-production-defaults -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install waybarrios/opencode-power-pack hf-cloud-sagemaker-production-defaults --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/waybarrios/opencode-power-pack.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/hf-cloud-sagemaker-production-defaults .claude/skills/hf-cloud-sagemaker-production-defaults && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "hf-cloud-sagemaker-production-defaults" agent skill from https://github.com/waybarrios/opencode-power-pack/tree/main/skills/hf-cloud-sagemaker-production-defaults into .claude/skills/hf-cloud-sagemaker-production-defaults/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "hf-cloud-sagemaker-production-defaults", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/waybarrios/opencode-power-pack/tree/main/skills/hf-cloud-sagemaker-production-defaultsType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add waybarrios/opencode-power-pack --skill hf-cloud-sagemaker-production-defaults -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install waybarrios/opencode-power-pack hf-cloud-sagemaker-production-defaults --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/waybarrios/opencode-power-pack.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/hf-cloud-sagemaker-production-defaults .agents/skills/hf-cloud-sagemaker-production-defaults && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "hf-cloud-sagemaker-production-defaults" agent skill from https://github.com/waybarrios/opencode-power-pack/tree/main/skills/hf-cloud-sagemaker-production-defaults into .agents/skills/hf-cloud-sagemaker-production-defaults/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "hf-cloud-sagemaker-production-defaults", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add waybarrios/opencode-power-pack --skill hf-cloud-sagemaker-production-defaults -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install waybarrios/opencode-power-pack hf-cloud-sagemaker-production-defaults --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/waybarrios/opencode-power-pack.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/hf-cloud-sagemaker-production-defaults .cursor/skills/hf-cloud-sagemaker-production-defaults && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "hf-cloud-sagemaker-production-defaults" agent skill from https://github.com/waybarrios/opencode-power-pack/tree/main/skills/hf-cloud-sagemaker-production-defaults into .cursor/skills/hf-cloud-sagemaker-production-defaults/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "hf-cloud-sagemaker-production-defaults", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/waybarrios/opencode-power-pack.git --path skills/hf-cloud-sagemaker-production-defaults--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add waybarrios/opencode-power-pack --skill hf-cloud-sagemaker-production-defaults -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install waybarrios/opencode-power-pack hf-cloud-sagemaker-production-defaults --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/waybarrios/opencode-power-pack.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/hf-cloud-sagemaker-production-defaults .gemini/skills/hf-cloud-sagemaker-production-defaults && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "hf-cloud-sagemaker-production-defaults" agent skill from https://github.com/waybarrios/opencode-power-pack/tree/main/skills/hf-cloud-sagemaker-production-defaults into .gemini/skills/hf-cloud-sagemaker-production-defaults/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "hf-cloud-sagemaker-production-defaults", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install waybarrios/opencode-power-pack hf-cloud-sagemaker-production-defaultsInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add waybarrios/opencode-power-pack --skill hf-cloud-sagemaker-production-defaults -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/waybarrios/opencode-power-pack.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/hf-cloud-sagemaker-production-defaults .github/skills/hf-cloud-sagemaker-production-defaults && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "hf-cloud-sagemaker-production-defaults" agent skill from https://github.com/waybarrios/opencode-power-pack/tree/main/skills/hf-cloud-sagemaker-production-defaults into .github/skills/hf-cloud-sagemaker-production-defaults/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "hf-cloud-sagemaker-production-defaults", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add waybarrios/opencode-power-pack --skill hf-cloud-sagemaker-production-defaults -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install waybarrios/opencode-power-pack hf-cloud-sagemaker-production-defaults --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/waybarrios/opencode-power-pack.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/hf-cloud-sagemaker-production-defaults .opencode/skills/hf-cloud-sagemaker-production-defaults && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "hf-cloud-sagemaker-production-defaults" agent skill from https://github.com/waybarrios/opencode-power-pack/tree/main/skills/hf-cloud-sagemaker-production-defaults into .opencode/skills/hf-cloud-sagemaker-production-defaults/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "hf-cloud-sagemaker-production-defaults", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
hf-cloud-sagemaker-production-defaultsImplement a production SageMaker endpoint with autoscaling, CloudWatch alarms, and tags.
Hf Cloud Sagemaker Production Defaults is an agent skill from waybarrios/opencode-power-pack. Implement a production SageMaker endpoint with autoscaling, CloudWatch alarms, and tags. Use after the serving image and IAM role are known; use the deployment planner first when architecture is undecided.
Its SKILL.md is about 4.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 8 other files, including scripts and reference files (for example `references/deployment-template.md`, `scripts/_common.py` and `scripts/deploy.py`).
It sits in AI & LLM Engineering, covering Deployment and LLM inference and serving. It works with Amazon SageMaker and vLLM. The repository describes itself as: 54 rigorous skills for Codex, OpenCode, and Pi: code review, security audit, feature development, frontend design, MCP tools, Hugging Face ML/training, and more. The licence is Apache-2.0.
5 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 9dccb6d. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 5 files in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
pythonawspython3From the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
aws.github.ioFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Hf Cloud Sagemaker Production Defaults loads about 4.6k tokens when it runs, and up to ~5.5k if it reads all its reference files. Until then it costs about 61 tokens; SKILL.md has 1,953 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from waybarrios/opencode-power-pack at commit 9dccb6d, republished under its Apache-2.0 licence (© waybarrios). 1,953 words, ~4,601 tokens.
.claude/skills/hf-cloud-sagemaker-production-defaults/SKILL.md (or your agent's skills folder). This skill also uses 6 other files; get the full folder from GitHub.The difference between a demo endpoint and one you can leave running is: it scales with traffic, it tells you when it breaks, and you can debug it later. This skill makes those three the default rather than optional extras.
By the time this skill runs, the planner has chosen a real-time endpoint, IAM has a usable role, and image-selection has resolved a container URI + AMI version. This skill turns those into an actual deployment.
For every endpoint, the skill creates these as a unit:
Data capture (logging requests/responses to S3) is off by default — useful for debugging but creates ongoing S3 costs the user didn't necessarily ask for. Enable with --enable-data-capture.
All resources get a consistent tag set including CreatedBy=agentic-deploy-skills for later cleanup.
Defaults and reasoning in references/deployment-template.md.
For a text-generation LLM (vLLM):
python scripts/deploy.py \
--model-name qwen3-medical \
--image-uri "$IMAGE_URI" \
--inference-ami-version "$AMI" \
--role-arn "$ROLE_ARN" \
--instance-type ml.g5.xlarge \
--region "$REGION" \
--env SM_VLLM_MODEL=Qwen/Qwen3-0.6B \
--env SM_VLLM_HOST=0.0.0.0 \
--env SM_VLLM_TRUST_REMOTE_CODE=true \
--env SM_VLLM_MAX_MODEL_LEN=4096For an embedding model (TEI, often on CPU):
python scripts/deploy.py \
--model-name bge-large-embeddings \
--image-uri "$IMAGE_URI" \
--role-arn "$ROLE_ARN" \
--instance-type ml.c6i.2xlarge \
--region "$REGION" \
--env HF_MODEL_ID=BAAI/bge-large-en-v1.5Note: TEI deployments do not need --inference-ami-version. That flag is vLLM-specific. TEI env vars are also simpler (HF_MODEL_ID instead of SM_VLLM_*, no host or trust-remote-code to configure).
Where each value comes from:
| Parameter | Source |
|---|---|
--image-uri | hf-cloud-serving-image-selection — agent reads from the AWS DLC catalog page |
--inference-ami-version | hf-cloud-serving-image-selection — required for vLLM tags containing cu130+ |
--role-arn | hf-cloud-sagemaker-iam-preflight (check_role.py) |
--region | hf-cloud-aws-context-discovery |
--instance-type | User input or planner recommendation |
--env | Model-specific; see hf-cloud-serving-image-selection for required SM_VLLM_* vars |
--model-s3-uri | Optional — S3 path to model artifacts; omit if loading from HF Hub |
The script creates resources in order with error handling, waits for InService (up to 30 min), surfaces failure reasons, registers autoscaling and alarms, and prints a summary including the teardown command. Outputs a JSON blob on stdout with endpoint/config/model names for downstream scripting.
The scripts ship with this skill. If the installed copy is missing the scripts/ directory (some harnesses copy only SKILL.md on install), fetch them from the source repo rather than re-implementing them from this description.
Cold-start expectation: when the model loads from HF Hub, the download happens inside the container after the endpoint starts — 5–15+ minutes to InService is normal, not a failure. deploy.py waits 30 minutes; if you write custom wait code, don't time out at 15. Pre-staging weights in S3 (--model-s3-uri) cuts this and removes the Hub dependency.
InService only means the container answered /ping. In MMS-based containers (HF Inference Toolkit) the Java front-end answers pings even while the Python worker crash-loops — an endpoint can be InService and serve nothing. Two checks, always:
One real invocation.
invoke_endpoint.py (below) with a minimal payload; require an HTTP 200 with a sane body.invoke-endpoint-async, poll the output URI for a few minutes (see "Invoking async endpoints"). A result object = success; an object at the failure URI, or nothing appearing, = broken.Scan the endpoint logs for worker-crash markers — catches the crash-loop case even when the smoke request merely times out:
aws logs filter-log-events \
--log-group-name /aws/sagemaker/Endpoints/<endpoint-name> \
--filter-pattern '?"Worker died" ?"Load model failed" ?"ImportError"' \
--region <region> --max-items 5Only report the deployment complete after both pass. If the log scan hits, surface the actual traceback from CloudWatch — not the InService status.
Once the endpoint is InService, test it with the bundled helper. It is cross-platform and BOM-safe — use it instead of hand-writing a payload file and calling invoke-endpoint directly:
# macOS / Linux
python3 scripts/invoke_endpoint.py \
--endpoint-name <endpoint-name> \
--payload '{"inputs": "Hello"}' \
--region "$REGION"# Windows (PowerShell)
python scripts\invoke_endpoint.py `
--endpoint-name <endpoint-name> `
--payload-file payload.json `
--region $REGIONIt accepts either --payload '<json>' (inline) or --payload-file <path>, validates JSON, writes the request body as plain UTF-8, invokes the endpoint, and prints the response body to stdout.
If you write the request payload yourself on Windows, do not use Set-Content -Encoding UTF8 — depending on the PowerShell version it prepends a UTF-8 byte-order mark (BOM). SageMaker's JSON parser rejects a BOM with a 400 ModelError:
Unexpected UTF-8 BOM (decode using utf-8-sig): line 1 column 1 (char 0)This is not a model, endpoint-health, or image problem — only the file encoding of the request body. invoke_endpoint.py avoids it entirely (it even strips a BOM from a --payload-file that already has one). If you must call the CLI directly, write the body as BOM-free UTF-8:
# BOM-free UTF-8 — use this
[System.IO.File]::WriteAllText((Resolve-Path "payload.json"), $json, [System.Text.UTF8Encoding]::new($false))
aws sagemaker-runtime invoke-endpoint `
--endpoint-name <endpoint-name> `
--content-type application/json `
--body fileb://payload.json `
--region $REGION `
response.jsonFallback: if any invocation fails with Unexpected UTF-8 BOM, rewrite the payload as BOM-free UTF-8 (or re-run via invoke_endpoint.py) and retry once before treating the endpoint or model as broken.
Generative rerankers (Qwen3-Reranker etc. — routed to the HuggingFace vLLM DLC by hf-cloud-serving-image-selection) are causal LMs scored by their first generated token, not chat models. Use the completions API with a raw prompt, not the messages/chat API: chat templating does not reliably honor chat_template_kwargs such as {"enable_thinking": false}, and a wrong template silently returns near-identical scores for every query–document pair instead of erroring.
Payload shape (Qwen3-Reranker's expected format — substitute {query} / {document}):
{
"prompt": "<|im_start|>system\nJudge whether the Document meets the requirements based on the Query and the Instruct provided. Note that the answer can only be \"yes\" or \"no\".<|im_end|>\n<|im_start|>user\n<Instruct>: Given a web search query, retrieve relevant passages that answer the query\n<Query>: {query}\n<Document>: {document}<|im_end|>\n<|im_start|>assistant\n<think>\n\n</think>\n\n",
"max_tokens": 1,
"temperature": 0,
"logprobs": 20
}The trailing <|im_start|>assistant\n<think>\n\n</think>\n\n suffix is load-bearing: it pre-fills an empty thinking block so the first generated token is the yes/no judgment. Score from the returned logprobs: P("yes") / (P("yes") + P("no")). Sanity check the endpoint with one relevant pair (expect >0.9) and one irrelevant pair (expect <0.05) — near-identical scores across pairs mean the prompt template is wrong, not that the model is broken.
The same rule generalizes: for any thinking-mode model where the prompt must be byte-exact, prefer the raw completions API over chat.
The agent reads the image URI from AWS's Deep Learning Containers catalog — pick the row that matches the model family (HuggingFace vLLM for LLMs, TEI for embeddings, etc.), substitute <region> with the deployment region, and pass to deploy.py --image-uri.
For vLLM images specifically (both huggingface-vllm and the AWS vllm fallback), also check the tag's CUDA version:
# Example: HuggingFace vLLM 0.21.0 from the catalog
IMAGE_URI="763104351884.dkr.ecr.eu-west-1.amazonaws.com/huggingface-vllm:0.21.0-transformers5.8.1-gpu-py312-cu130-ubuntu22.04"
# cu130 tag → must pass --inference-ami-version
python deploy.py --image-uri "$IMAGE_URI" \
--inference-ami-version al2-ami-sagemaker-inference-gpu-3-1 \
...For tags with cu129 or lower, omit --inference-ami-version. See hf-cloud-serving-image-selection for the full vLLM AMI lookup table and the env-var requirements for each image family.
For long-running inferences (>60s), large payloads, or workloads that are bursty/sparse enough to benefit from scale-to-zero, use deploy_async.py instead of deploy.py. Async genuinely supports MinCapacity=0 — real-time autoscaling can't.
python scripts/deploy_async.py \
--model-name flux-text-to-image \
--image-uri "$IMAGE_URI" \
--role-arn "$ROLE_ARN" \
--instance-type ml.g5.2xlarge \
--region "$REGION" \
--output-s3-uri s3://my-bucket/async-output/ \
--env HF_MODEL_ID=black-forest-labs/FLUX.1-devRequired extras over deploy.py:
--output-s3-uri — where async results land (results are not returned synchronously)Optional async-specific flags:
--failure-s3-uri — separate path for failed invocations--success-sns-topic, --error-sns-topic — get notified when async results are ready or fail--min-capacity 0 (the default) — scale to zero between batches--backlog-per-instance-target N — target queue depth per instance (default 5)--max-concurrent-invocations-per-instance N — default 4The async script registers two autoscaling policies on the variant:
ApproximateBacklogSizePerInstance — handles ongoing scaling between min and maxHasBacklogWithoutCapacity CloudWatch alarm — handles 0→1 wake-from-zeroBoth are needed. Target-tracking alone cannot transition from zero (it can't divide by zero instances), so without the step policy the endpoint comes up, scales to zero after the first batch, and never wakes again. The script wires this up automatically.
The script creates three CloudWatch alarms:
ApproximateBacklogSize > 50 — queue is building faster than capacity can drain itInvocationsFailed > 5 — repeated processing failuresHasBacklogWithoutCapacity — drives the wake-from-zero policy (not a notification alarm; its action is the step-scaling policy, not the SNS topic)If you pass --sns-alarm-topic <arn>, the first two notify on that topic. The wake alarm always points at the step policy.
Async endpoints aren't called synchronously. You upload the input to S3, call invoke-endpoint-async with the S3 input location, and SageMaker writes the result to your --output-s3-uri when done:
# Upload your input first
aws s3 cp input.json s3://my-input-bucket/job1/input.json
# Invoke
aws sagemaker-runtime invoke-endpoint-async \
--endpoint-name <endpoint-name> \
--input-location s3://my-input-bucket/job1/input.json \
--content-type application/json \
--region <region>
# Poll for the result at your output URI
aws s3 cp s3://my-bucket/async-output/<inference-id>.out result.jsonThe same UTF-8 BOM caveat applies to the input.json you upload (see "The UTF-8 BOM gotcha" above) — if you build it on Windows, write it as BOM-free UTF-8 or the container's JSON parser will reject it.
Teardown works the same as real-time: python3 scripts/teardown.py <endpoint-name> (the teardown script discovers policies and alarms by name prefix, so it handles both deployment modes).
| Setting | Default | Override |
|---|---|---|
| Initial instance count | 1 | --initial-instance-count |
| Autoscaling min / max | 1 / 4 | --min-capacity, --max-capacity |
| Autoscaling target | 20 invocations/min/instance | --target-invocations-per-instance |
| Data capture | disabled (opt-in) | --enable-data-capture |
| CloudWatch alarms | 3 alarms | --no-alarms |
| SNS notification | none (alarms created but won't notify) | --sns-alarm-topic <arn> |
| Environment tag | dev | --environment |
| InferenceAmiVersion | none (SageMaker default) | --inference-ami-version (REQUIRED for vLLM CUDA 13+) |
Not defaulted (user-specific input needed): VPC config, KMS key, multi-variant, async inference.
The default --target-invocations-per-instance 20 is conservative and tuned for LLM workloads where each request takes 1–5 seconds. For embedding deployments (TEI), each request is much faster (typically <100ms on CPU, <20ms on GPU), so a single instance can handle far more throughput. For embedding deployments, raise the target to 100–500 depending on instance and model size. The default of 20 will trigger autoscaling far too aggressively for embeddings and waste money.
A rule of thumb: target value ≈ 60 / (typical request latency in seconds). LLM at 3s latency → target 20. Embedding at 100ms → target 600. Generative rerankers sit in between — they generate a single token per request, so ~40–100 is a reasonable target.
If the user enables data capture, the execution role needs S3 write access to the capture prefix. The default URI (s3://sagemaker-<region>-<account>/<endpoint>/data-capture/) is typically a different bucket than the model artifact bucket. If hf-cloud-sagemaker-iam-preflight scoped the inline policy narrowly to just the model bucket, capture writes fail silently — endpoint keeps serving but no data appears.
If the user reports "data capture isn't showing up", check the role's S3 access. Either widen the inline policy or pass --data-capture-s3-uri pointing to a bucket the role can write.
python3 scripts/teardown.py <endpoint-name> <region> # macOS / Linux
python scripts\teardown.py <endpoint-name> <region> # WindowsDeletes in safe order: alarms → autoscaling → endpoint (stops billing) → endpoint config → model. Idempotent.
Does not delete: the IAM execution role (might be shared), data capture S3 objects (user might want to keep), SNS topic, original model artifacts.
Always tell the user about the teardown command after the deployment summary. Users forget; endpoints accrue cost.
CannotStartContainerError + no CloudWatch logs ever created — the InferenceAmiVersion problem. If the image tag contains cu130 or later and you didn't pass --inference-ami-version al2-ami-sagemaker-inference-gpu-3-1, this is the cause. See hf-cloud-serving-image-selection. Do NOT chase images, IAM roles, env vars, or instance types — the failure signature is identical for many other things but the cause here is the AMI.
"Failed to pass ping health check" — the container did start and produced logs, but /ping isn't responding. Check CloudWatch at /aws/sagemaker/Endpoints/<endpoint-name>. Usually: wrong image for model architecture, missing HF token, or OOM.
"Container failed to start" (with logs present) — entrypoint ran, then exited. Check CloudWatch. Common: missing required env vars (SM_VLLM_MODEL, SM_VLLM_HOST, SM_VLLM_TRUST_REMOTE_CODE), wrong ModelDataUrl format, unreadable model artifacts.
ResourceLimitExceeded — no quota for the instance type in this region. Request increase or pick a different type (the planner should have checked quotas up front — see hf-cloud-sagemaker-deployment-planner).
ImportError: libtorch_cuda.so: undefined symbol: ncclCommResume in CloudWatch logs — known packaging defect in huggingface-pytorch-inference GPU images (see "Known-broken images" in hf-cloud-serving-image-selection). Inside the container, so no env var, AMI, instance type, or sibling tag fixes it. Switch to DJL Inference.
InService, but invocations time out / async outputs never appear — dead Python worker behind a live MMS front-end. Run the log scan from "InService is not success" above; the traceback in CloudWatch is the real error.
403 Forbidden downloading weights from HF Hub during startup — the container's bundled huggingface_hub predates HF's XET CDN auth. Add --env HF_HUB_ENABLE_HF_TRANSFER=0, or pre-stage the weights in S3. Note: this can mask a deeper failure (the worker may still crash after the download succeeds) — re-check logs after fixing it.
Diagnostic rule: when failures look identical across multiple configurations (different images, roles, instance types) and no logs are ever produced, the cause is almost always below the container — host AMI, networking, account-level — not the deployment config. Stop iterating on config; check the AMI version and account state.
Don't retry blindly. The script prints the specific FailureReason from describe-endpoint — fix the root cause before retrying.
© waybarrios, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 6 other files (scripts, references) in skills/hf-cloud-sagemaker-production-defaults of waybarrios/opencode-power-pack.
Open the folder on GitHubat commit 9dccb6d
Hf Cloud Sagemaker Production Defaults next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Hf Cloud Sagemaker Production Defaults this skillwaybarrios/opencode-power-pack | 534 | — | ~4.6k | Automated safety check: Pass | Apache-2.0 | |
| SageMaker Serving Image Selectionhuggingface/skills | 11k | 1 repos | ~4.6k | Automated safety check: Pass | Apache-2.0 | |
| SageMaker Production Defaultshuggingface/skills | 11k | 1 repos | ~6.9k | Automated safety check: Pass | Apache-2.0 | |
| vLLM Model ServingOrchestra-Research/AI-Research-SKILLs | 13k | 5 repos | ~2.3k | Automated safety check: Pass | MIT | |
| Vllm Deploy K8svllm-project/vllm-skills | 102 | — | ~2k | Automated safety check: Pass | Apache-2.0 | |
| SageMaker Deployment Plannerhuggingface/skills | 11k | 1 repos | ~2.1k | Automated safety check: Pass | Apache-2.0 |
huggingface/skills
Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.
huggingface/skills
Deploys SageMaker endpoints with autoscaling, CloudWatch alarms and tags on by default, using scripts for real-time, scale-to-zero and async setups.
Orchestra-Research/AI-Research-SKILLs
Deploys LLMs with vLLM for high-throughput serving, covering the OpenAI-compatible server, offline batch inference, monitoring and a Docker rollout.
vllm-project/vllm-skills
Deploy vLLM to Kubernetes (K8s) with GPU support, health probes, and OpenAI-compatible API endpoint.
huggingface/skills
Entry point for hosting a model on Amazon SageMaker: asks a few questions, picks a deployment pathway and hands off to the specialist skills.
aws/agent-toolkit-for-aws
Selects, deploys, and customizes AI models on Amazon SageMaker.
waybarrios/opencode-power-pack
Verify or select a SageMaker execution role before creating models, endpoints, or training jobs.
waybarrios/opencode-power-pack
Train or fine-tune language models with TRL or Unsloth on Hugging Face Jobs, including SFT, DPO, GRPO, reward models, and GGUF conversion.
waybarrios/opencode-power-pack
Train object-detection, image-classification, or SAM segmentation models on Hugging Face Jobs.
waybarrios/opencode-power-pack
Run CodeQL database creation and security queries, add data-extension models, or process CodeQL SARIF.
waybarrios/opencode-power-pack
Run Semgrep static analysis across a codebase, optionally using Semgrep Pro for cross-file taint analysis.
waybarrios/opencode-power-pack
Detects fail-open insecure defaults (hardcoded secrets, weak auth, permissive security) that allow apps to run insecurely in production.
Works with
Categories
Implement a production SageMaker endpoint with autoscaling, CloudWatch alarms, and tags. Hf Cloud Sagemaker Production Defaults is an agent skill from waybarrios/opencode-power-pack. Implement a production SageMaker endpoint with autoscaling, CloudWatch alarms, and tags.
Hf Cloud Sagemaker Production Defaults fits situations like: tasks that involve Deployment; tasks that involve LLM inference and serving.
Run `npx skills add waybarrios/opencode-power-pack --skill hf-cloud-sagemaker-production-defaults -a claude-code`. Or copy the skill folder (skills/hf-cloud-sagemaker-production-defaults in waybarrios/opencode-power-pack) into .claude/skills/hf-cloud-sagemaker-production-defaults in your project. Claude Code loads it when a task matches its description.
Run `npx skills add waybarrios/opencode-power-pack --skill hf-cloud-sagemaker-production-defaults -a codex`. Or copy the skill folder (skills/hf-cloud-sagemaker-production-defaults in waybarrios/opencode-power-pack) into .agents/skills/hf-cloud-sagemaker-production-defaults in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add waybarrios/opencode-power-pack --skill hf-cloud-sagemaker-production-defaults -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/hf-cloud-sagemaker-production-defaults, .gemini/skills/hf-cloud-sagemaker-production-defaults, .github/skills/hf-cloud-sagemaker-production-defaults and .opencode/skills/hf-cloud-sagemaker-production-defaults in your project.
Going by SKILL.md and its folder, Hf Cloud Sagemaker Production Defaults needs Python for the scripts in its folder and the command-line tools its instructions call (python, aws and python3). Our summary lists: Python 3.
SKILL.md names 1 domain. As links in the text: aws.github.io. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Hf Cloud Sagemaker Production Defaults is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.6k tokens (SKILL.md is roughly 18k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 916 tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Hf Cloud Sagemaker Production Defaults: SageMaker Serving Image Selection (huggingface/skills, 11k stars), SageMaker Production Defaults (huggingface/skills, 11k stars), vLLM Model Serving (Orchestra-Research/AI-Research-SKILLs, 13k stars) and Vllm Deploy K8s (vllm-project/vllm-skills, 102 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
waybarrios (a GitHub user) maintains it in waybarrios/opencode-power-pack, which has 534 GitHub stars. The repository holds 32 skills in this directory. The repository was last updated on October 6, 2026.
Source: waybarrios/opencode-power-pack on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.