Search
Amazon SageMaker
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Finds or validates a usable SageMaker execution role before deploying or training, so scripts do not try to create IAM roles they lack permission to create. | huggingface/ | 11k | 1 repo | ~1.8k | Automated safety check: Pass | Apache-2.0 | today |
| 2 | Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones. | huggingface/ | 11k | 1 repo | ~4.6k | Automated safety check: Pass | Apache-2.0 | today |
| 3 | Sets up an isolated Python environment with a supported interpreter and current boto3 before any SageMaker deployment, training or AWS automation code runs. | huggingface/ | 11k | 2 repos | ~1.7k | Automated safety check: Pass | Apache-2.0 | today |
| 4 | Validates dataset formatting and quality for SageMaker model fine-tuning (SFT, DPO, or RLVR). | awslabs/ | 915 | 1 repo | ~1.3k | Automated safety check: Pass | Apache-2.0 | today |
| 5 | Generates code that transforms datasets between ML schemas for model training or evaluation. | awslabs/ | 915 | 1 repo | ~3.5k | Automated safety check: Pass | Apache-2.0 | today |
| 6 | Deploys SageMaker endpoints with autoscaling, CloudWatch alarms and tags on by default, using scripts for real-time, scale-to-zero and async setups. | huggingface/ | 11k | 1 repo | ~6.9k | Automated safety check: Pass | Apache-2.0 | today |
| 7 | Verify or select a SageMaker execution role before creating models, endpoints, or training jobs. | waybarrios/ | 533 | — | ~1.6k | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 8 | Selects a fine-tuning technique (SFT, DPO, RLVR, or RLAIF) for the user's use case and validates it against the selected model's available recipes. | awslabs/ | 915 | — | ~604 | Automated safety check: Pass | Apache-2.0 | today |
| 9 | Generate comprehensive issue reports from HyperPod clusters (EKS and Slurm) by collecting diagnostic logs and configurations for troubleshooting and AWS Support cases. | awslabs/ | 915 | — | ~890 | Automated safety check: Pass | Apache-2.0 | today |
| 10 | A skill your agent uses for GPU training or inference clusters on SageMaker HyperPod (Slurm or EKS), ParallelCluster, or self-managed EC2/EKS GPU instances. | aws/ | 102 | — | ~5.4k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 11 | Diagnose performance issues on Amazon SageMaker HyperPod clusters — uneven NCCL bandwidth across nodes and poor filesystem throughput. | awslabs/ | 915 | — | ~4.1k | Automated safety check: Pass | Apache-2.0 | today |
| 12 | Diagnostic-only skill for Slurm scheduler and node-daemon issues on Amazon SageMaker HyperPod Slurm clusters. | awslabs/ | 915 | — | ~3.3k | Automated safety check: Pass | Apache-2.0 | today |
| 13 | Entry point for hosting a model on Amazon SageMaker: asks a few questions, picks a deployment pathway and hands off to the specialist skills. | huggingface/ | 11k | 1 repo | ~2.1k | Automated safety check: Pass | Apache-2.0 | today |
| 14 | Remote command execution and file transfer on SageMaker HyperPod cluster nodes via AWS Systems Manager (SSM). | awslabs/ | 915 | — | ~1.3k | Automated safety check: Notes | Apache-2.0 | today |
| 15 | Uses Meta's LlamaGuard moderation model to screen prompts and model replies against six safety categories, with vLLM, FastAPI and NeMo Guardrails setups. | Orchestra-Research/ | 13k | 2 repos | ~2.3k | Automated safety check: Pass | MIT | 3 mo ago |
| 16 | Selects a base model for the user's use case by querying SageMaker Hub. | awslabs/ | 915 | — | ~844 | Automated safety check: Pass | Apache-2.0 | today |
| 17 | Deep expertise in ML/CV model selection, training pipelines, and inference architecture. | alirezarezvani/ | 117 | — | ~3.1k | Automated safety check: Pass | MIT | 9 mo ago |
| 18 | Discover the user's local AWS context (active profile, region, account ID, caller identity) at the start of any AWS task. | huggingface/ | 11k | 2 repos | ~989 | Automated safety check: Warn | Apache-2.0 | today |
| 19 | Amazon SageMaker AI Operational Review. An agent skill from aws/tools-for-devops-agent. | aws/ | 102 | — | ~3.9k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 20 | Implement a production SageMaker endpoint with autoscaling, CloudWatch alarms, and tags. | waybarrios/ | 533 | — | ~4.6k | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 21 | Check and compare software component versions on SageMaker HyperPod cluster nodes - NVIDIA drivers, CUDA toolkit, cuDNN, NCCL, EFA, AWS OFI NCCL, GDRCopy, MPI, Neuron SDK (Trainium/Inferentia)… | awslabs/ | 915 | — | ~910 | Automated safety check: Pass | Apache-2.0 | today |
| 22 | Select and verify the current region-specific serving container URI for a SageMaker model deployment. | waybarrios/ | 533 | — | ~4.3k | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 23 | Generates code that fine-tunes a base model using SageMaker serverless training jobs. | awslabs/ | 915 | — | ~2.3k | Automated safety check: Pass | Apache-2.0 | today |
| 24 | Generates python code that evaluates SageMaker models. An agent skill from awslabs/agent-plugins. | awslabs/ | 915 | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | today |
| 25 | Diagnose NCCL failures and adjacent training-pod failures on HyperPod GPU clusters (EKS or Slurm) — training hangs, AllReduce / collective-op timeouts, EFA or libfabric errors, rendezvous failures… | awslabs/ | 915 | — | ~3.4k | Automated safety check: Pass | Apache-2.0 | today |
| 26 | Generates code that deploys fine-tuned models from SageMaker Serverless Model Customization to SageMaker endpoints or Bedrock. | awslabs/ | 915 | — | ~1.5k | Automated safety check: Pass | Apache-2.0 | today |
| 27 | Validates the user's environment for SageMaker AI operations — checks SDK version, AWS region, and execution role. | awslabs/ | 915 | — | ~248 | Automated safety check: Pass | Apache-2.0 | today |
| 28 | Plan and coordinate a model deployment to Amazon SageMaker, including serving stack and real-time versus async inference. | waybarrios/ | 533 | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 29 | Runs SQL analytics on SageMaker Catalog asset metadata tables exported as Apache Iceberg in S3 Tables. | aws/ | 2.8k | — | ~2.6k | Automated safety check: Pass | Apache-2.0 | today |
| 30 | Selects, deploys, and customizes AI models on Amazon SageMaker. | aws/ | 2.8k | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | today |
| 31 | A skill your agent uses when diagnosing IAM and access failures for Bedrock and SageMaker. | aws/ | 102 | — | ~4.9k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 32 | Cost guardrail for AWS DevOps Agent that covers ALL AWS services and native agent tools. | aws/ | 102 | — | ~4.5k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 33 | Deploy sagemaker endpoint deployer operations. An agent skill from jeremylongshore/tons-of-skills-marketplace. | jeremylongshore/ | 2.8k | — | ~586 | Automated safety check: Pass | MIT | today |
| 34 | A full ML pipeline where an agent team collaborates to perform data preparation, model design, training, evaluation, and deployment readiness. | revfactory/ | 1.3k | — | ~1.9k | Automated safety check: Pass | Apache-2.0 | 6 mo ago |