Agent skill

Hf Cloud Sagemaker Deployment Planner

by waybarrios in waybarrios/opencode-power-pack

Plan and coordinate a model deployment to Amazon SageMaker, including serving stack and real-time versus async inference.

Apache-2.0Auto-check passedDevOps & Cloud

Install Hf Cloud Sagemaker Deployment Planner

skills CLI
$ npx skills add waybarrios/opencode-power-pack --skill hf-cloud-sagemaker-deployment-planner -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install waybarrios/opencode-power-pack hf-cloud-sagemaker-deployment-planner --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/waybarrios/opencode-power-pack.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/hf-cloud-sagemaker-deployment-planner .claude/skills/hf-cloud-sagemaker-deployment-planner && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
hf-cloud-sagemaker-deployment-planner
GitHub stars
533
Token cost
~1.7k tokens
SKILL.md length
892 words
Files
1
Skills in repo
32
Repo updated
First seen
Licence
Apache-2.0

At a glance

Plan and coordinate a model deployment to Amazon SageMaker, including serving stack and real-time versus async inference.

  • Works in 6 steps: Discovery — what is being deployed and… → Pathway selection — real-time /… → Context preflight —… → …
  • Tasks that involve Deployment
  • SKILL.md covers Workflow phases, Discovery: ask only what you…, Pathway selection and Instance selection: check…, plus 1 more section
  • Calls aws

What it does

Hf Cloud Sagemaker Deployment Planner is an agent skill from waybarrios/opencode-power-pack. Plan and coordinate a model deployment to Amazon SageMaker, including serving stack and real-time versus async inference. Use as the entry point for SageMaker hosting requests, before image, IAM, and endpoint implementation skills.

Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in DevOps & Cloud, covering Deployment. It works with Amazon SageMaker. The repository describes itself as: 54 rigorous skills for Codex, OpenCode, and Pi: code review, security audit, feature development, frontend design, MCP tools, Hugging Face ML/training, and more. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Deployment

Example prompts

  • “/hf-cloud-sagemaker-deployment-planner”

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Discovery — what is being deployed and what are the constraints (this skill)
  2. Pathway selection — real-time / serverless / async / batch / Bedrock CMI (this skill)
  3. Context preflight — hf-cloud-aws-context-discovery, then hf-cloud-python-env-setup
  4. IAM preflight — hf-cloud-sagemaker-iam-preflight
  5. Image selection — hf-cloud-serving-image-selection
  6. Deployment — hf-cloud-sagemaker-production-defaults

What it can do on your machine

Read from SKILL.md and the folder at commit 9dccb6d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • aws

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use aws, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Hf Cloud Sagemaker Deployment Planner loads about 1.7k tokens when it runs. Until then it costs about 67 tokens; SKILL.md has 892 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~67
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from waybarrios/opencode-power-pack at commit 9dccb6d, republished under its Apache-2.0 licence (© waybarrios). 892 words, ~1,703 tokens.

Download SKILL.mdSave it as .claude/skills/hf-cloud-sagemaker-deployment-planner/SKILL.md (or your agent's skills folder).
name
hf-cloud-sagemaker-deployment-planner
description
Plan and coordinate a model deployment to Amazon SageMaker, including serving stack and real-time versus async inference. Use as the entry point for SageMaker hosting requests, before image, IAM, and endpoint implementation skills.
license
Apache-2.0 (modified; see UPSTREAMS.json)

SageMaker Deployment Planner

You are helping a user deploy a model to Amazon SageMaker. Most users invoking this skill want the model deployed with reasonable defaults, in as few questions as possible. Ask only what you need, recommend a pathway honestly, and hand off to the specialized skills.

Workflow phases

  1. Discovery — what is being deployed and what are the constraints (this skill)
  2. Pathway selection — real-time / serverless / async / batch / Bedrock CMI (this skill)
  3. Context preflight — hf-cloud-aws-context-discovery, then hf-cloud-python-env-setup
  4. IAM preflight — hf-cloud-sagemaker-iam-preflight
  5. Image selection — hf-cloud-serving-image-selection
  6. Deployment — hf-cloud-sagemaker-production-defaults

Phases 1–2 are this skill's job. The others activate when their patterns match.

Discovery: ask only what you need

You will eventually need to know:

  • What model: HuggingFace ID, S3 path to artifacts, or model name. If the user is vague ("the model I fine-tuned"), ask for the artifact location.
  • Model type: text-generation LLM, embedding/reranker, or other (classifier, NER, etc.). This determines the serving stack — usually inferable from the model name (anything ending in -embed-*, starting with BAAI/bge-, sentence-transformers/* etc. is embeddings; chat/instruct models are LLMs). Only ask if it's genuinely ambiguous.
  • Traffic shape: roughly how often will this be called?
  • Latency tolerance: interactive, near-real-time, or async?
  • Cost sensitivity: ask only if the user signals it or the traffic pattern is ambiguous.

Region comes from hf-cloud-aws-context-discovery — don't ask unless the user volunteers it.

Do not front-load all of these. A common minimal set is just: what model, and roughly how often will it be called? The model name usually settles the model-type question. That alone is often enough to narrow the pathway to two candidates. If the user already told you something, don't ask again.

Pathway selection

PathwayWhen it fitsWhen it does not
Real-time endpointSteady traffic, sub-second to few-second latency, always-onVery spiky or very sparse traffic (wastes money on idle)
Serverless inferenceSpiky/intermittent, tolerates cold starts (~10s+), simpler modelsLLMs above a few B params (memory/cold-start limits), strict SLAs
Async inferenceLong inference (>60s), large payloads, queue-friendlyInteractive synchronous calls
Batch transformOffline scoring over a datasetAnything online or interactive
Bedrock Custom Model ImportWants Bedrock-compatible API, supported base family, weights onlyCustom inference logic, unsupported architectures

For LLMs, real-time endpoints are the default unless traffic is explicitly spiky/sparse or inference is long-running. Serverless looks attractive for "low traffic" cases but most LLMs exceed its memory limits.

For embeddings, real-time is again the default — but CPU instances are usually the right choice (much cheaper, fast enough for most embedding workloads). Don't reflexively recommend GPU instances for embedding models; ask hf-cloud-serving-image-selection to consider CPU variants if the model is small (<1B params) and traffic is moderate.

For text-to-image, video generation, or other long-inference workloads (>30s per request) where traffic is also bursty: async inference is the right answer. It supports genuine scale-to-zero between batches and queues requests via S3, so you don't pay for idle GPU. hf-cloud-sagemaker-production-defaults has a dedicated deploy_async.py for this.

Real-time and async are the two scripted pathways. Serverless, batch transform, and Bedrock Custom Model Import are not currently scripted — for those, hand the user off with a brief explanation rather than trying to deploy them through this workflow.

If two pathways are both reasonable, say so in one sentence each and pick one. Don't bury the recommendation in options.

Show full SKILL.md (344 more words)Show less

Instance selection: check quota before recommending

Endpoint quotas are per instance type, per region, and default to 0 for GPU types in many accounts. Recommending an instance the account can't launch wastes a full deploy cycle on ResourceLimitExceeded. Check first:

bash
aws service-quotas list-service-quotas --service-code sagemaker --region <region> \
    --query "Quotas[?contains(QuotaName, 'for endpoint usage') && Value > \`0\`].[QuotaName, Value]" \
    --output table

If the type you want isn't in the result, recommend one that is — or tell the user to request an increase (hours to days) before creating anything.

GPU family notes for the common 24 GB tier:

  • ml.g5.* (A10G) and ml.g6.* (L4) both work with current vLLM images when the gpu-3-1 AMI is set (see hf-cloud-serving-image-selection). g6 is the newer generation and slightly cheaper per hour; g5 has roughly double the memory bandwidth, which usually means better LLM token throughput. Pick whichever has quota; when both do, either is defensible — g5 for throughput, g6 for cost.
  • ml.g6e.* (L40S, 48 GB) when the model doesn't fit in 24 GB.

Once you have enough to recommend, state it plainly:

Based on what you've told me, I'd recommend a real-time endpoint on ml.g5.xlarge. The model is small enough that this is cost-effective, and your traffic pattern is steady enough that you won't be paying for idle. Alternative: serverless would be cheaper if traffic dries up for hours at a time, but Qwen3-0.6B is at the edge of serverless memory limits and cold starts would be 15–30s. Want me to proceed with the real-time endpoint?

Then wait for confirmation. The user should know what they're about to spend money on before you create anything.

The plan lives in the conversation — don't generate plan.yaml or similar artifacts unless explicitly asked.

Style

  • Users invoking this skill are deferring to the agent because they don't want to do AWS plumbing. Match that energy: efficient, not exhaustive.
  • One round of clarifying questions is usually enough. Three rounds is interrogation.
  • When you don't know something specific (current image URI, SDK API surface, quotas), check it rather than guess. Other skills handle the "how to check" details.
  • If the user pushes back on a recommendation, accept it. They know their constraints better than you do.

© waybarrios, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/hf-cloud-sagemaker-deployment-planner of waybarrios/opencode-power-pack.

Open the folder on GitHubat commit 9dccb6d

Compare with similar skills

Hf Cloud Sagemaker Deployment Planner next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Hf Cloud Sagemaker Deployment Planner compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Hf Cloud Sagemaker Deployment Planner this skillwaybarrios/opencode-power-pack533—~1.7kAutomated safety check: PassApache-2.0
SageMaker Production Defaultshuggingface/skills11k1 repos~6.9kAutomated safety check: PassApache-2.0
Sagemaker Endpoint Deployerjeremylongshore/tons-of-skills-marketplace2.8k—~586Automated safety check: PassMIT
SageMaker Serving Image Selectionhuggingface/skills11k1 repos~4.6kAutomated safety check: PassApache-2.0
Python Environment Setup for SageMakerhuggingface/skills11k2 repos~1.7kAutomated safety check: PassApache-2.0
SageMaker Deployment Plannerhuggingface/skills11k1 repos~2.1kAutomated safety check: PassApache-2.0

Similar skills

  • Official

    Deploys SageMaker endpoints with autoscaling, CloudWatch alarms and tags on by default, using scripts for real-time, scale-to-zero and async setups.

    11k GitHub starsUsed in 1 repo~6.9k tokens
    DevOps & CloudAuto-check passed
  • Sagemaker Endpoint Deployer

    jeremylongshore/tons-of-skills-marketplace

    Deploy sagemaker endpoint deployer operations. An agent skill from jeremylongshore/tons-of-skills-marketplace.

    2.8k GitHub stars~586 tokensUpdated today
    DevOps & CloudAuto-check passed
  • Official

    Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.

    11k GitHub starsUsed in 1 repo~4.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Sets up an isolated Python environment with a supported interpreter and current boto3 before any SageMaker deployment, training or AWS automation code runs.

    11k GitHub starsUsed in 2 repos~1.7k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Entry point for hosting a model on Amazon SageMaker: asks a few questions, picks a deployment pathway and hands off to the specialist skills.

    11k GitHub starsUsed in 1 repo~2.1k tokens
    AI & LLM EngineeringAuto-check passed
  • Model Deployment

    awslabs/agent-plugins

    Official

    Generates code that deploys fine-tuned models from SageMaker Serverless Model Customization to SageMaker endpoints or Bedrock.

    915 GitHub stars~1.5k tokensUpdated yesterday
    Backend & APIsAuto-check passed

More from waybarrios/opencode-power-pack

All 32 skills in this repo
  • Hf Cloud Sagemaker Iam Preflight

    waybarrios/opencode-power-pack

    Verify or select a SageMaker execution role before creating models, endpoints, or training jobs.

    533 GitHub stars~1.6k tokensUpdated 3 days ago
    Auto-check passed
  • Huggingface LLM Trainer

    waybarrios/opencode-power-pack

    Train or fine-tune language models with TRL or Unsloth on Hugging Face Jobs, including SFT, DPO, GRPO, reward models, and GGUF conversion.

    533 GitHub stars~3k tokensUpdated 3 days ago
    Auto-check passed
  • Huggingface Vision Trainer

    waybarrios/opencode-power-pack

    Train object-detection, image-classification, or SAM segmentation models on Hugging Face Jobs.

    533 GitHub stars~2.7k tokensUpdated 3 days ago
    Auto-check passed
  • Codeql

    waybarrios/opencode-power-pack

    Run CodeQL database creation and security queries, add data-extension models, or process CodeQL SARIF.

    533 GitHub starsUsed in 2 repos~3.7k tokens
    Auto-check passed
  • Semgrep

    waybarrios/opencode-power-pack

    Run Semgrep static analysis across a codebase, optionally using Semgrep Pro for cross-file taint analysis.

    533 GitHub stars~2.4k tokensUpdated 3 days ago
    Auto-check passed
  • Insecure Defaults

    waybarrios/opencode-power-pack

    Detects fail-open insecure defaults (hardcoded secrets, weak auth, permissive security) that allow apps to run insecurely in production.

    533 GitHub starsUsed in 2 repos~1.3k tokens
    Auto-check passed

Questions about Hf Cloud Sagemaker Deployment Planner

What does Hf Cloud Sagemaker Deployment Planner do?

Plan and coordinate a model deployment to Amazon SageMaker, including serving stack and real-time versus async inference. Hf Cloud Sagemaker Deployment Planner is an agent skill from waybarrios/opencode-power-pack. Plan and coordinate a model deployment to Amazon SageMaker, including serving stack and real-time versus async inference.

When should I use Hf Cloud Sagemaker Deployment Planner?

Hf Cloud Sagemaker Deployment Planner fits situations like: tasks that involve Deployment.

How do I install Hf Cloud Sagemaker Deployment Planner in Claude Code?

Run `npx skills add waybarrios/opencode-power-pack --skill hf-cloud-sagemaker-deployment-planner -a claude-code`. Or copy the skill folder (skills/hf-cloud-sagemaker-deployment-planner in waybarrios/opencode-power-pack) into .claude/skills/hf-cloud-sagemaker-deployment-planner in your project. Claude Code loads it when a task matches its description.

How do I install Hf Cloud Sagemaker Deployment Planner in Codex?

Run `npx skills add waybarrios/opencode-power-pack --skill hf-cloud-sagemaker-deployment-planner -a codex`. Or copy the skill folder (skills/hf-cloud-sagemaker-deployment-planner in waybarrios/opencode-power-pack) into .agents/skills/hf-cloud-sagemaker-deployment-planner in your project. Codex loads it when a task matches its description.

Can I use Hf Cloud Sagemaker Deployment Planner in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add waybarrios/opencode-power-pack --skill hf-cloud-sagemaker-deployment-planner -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/hf-cloud-sagemaker-deployment-planner, .gemini/skills/hf-cloud-sagemaker-deployment-planner, .github/skills/hf-cloud-sagemaker-deployment-planner and .opencode/skills/hf-cloud-sagemaker-deployment-planner in your project.

What does Hf Cloud Sagemaker Deployment Planner need to run?

Going by SKILL.md and its folder, Hf Cloud Sagemaker Deployment Planner needs the command-line tools its instructions call (aws).

Does Hf Cloud Sagemaker Deployment Planner access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Hf Cloud Sagemaker Deployment Planner safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Hf Cloud Sagemaker Deployment Planner use?

Hf Cloud Sagemaker Deployment Planner is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Hf Cloud Sagemaker Deployment Planner use?

About 1.7k tokens (SKILL.md is roughly 6.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Hf Cloud Sagemaker Deployment Planner?

Skills that share tags, products or a category with Hf Cloud Sagemaker Deployment Planner: SageMaker Production Defaults (huggingface/skills, 11k stars), Sagemaker Endpoint Deployer (jeremylongshore/tons-of-skills-marketplace, 2.8k stars), SageMaker Serving Image Selection (huggingface/skills, 11k stars) and Python Environment Setup for SageMaker (huggingface/skills, 11k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Hf Cloud Sagemaker Deployment Planner?

waybarrios (a GitHub user) maintains it in waybarrios/opencode-power-pack, which has 533 GitHub stars. The repository holds 32 skills in this directory. The repository was last updated on October 6, 2026.

Source: waybarrios/opencode-power-pack on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.