Topic · DevOps & Cloud
Best MLOps skills for Claude Code, Codex and other agents.
- skills
- 101
- official
- 11
MLOps skills, ranked
Ranked by score. Sort bymost stars,trending,newest,recently updated
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Finds or validates a usable SageMaker execution role before deploying or training, so scripts do not try to create IAM roles they lack permission to create. | huggingface/ | 11k | 1 repo | ~1.8k | Automated safety check: Pass | Apache-2.0 | 6 days ago |
| 2 | 2.Py Comprehensive Python programming reference covering syntax, concurrency, networking, databases, ML/LLM development, and HPC. | crazyguitar/ | 8.2k | — | ~886 | Automated safety check: Pass | MIT | yesterday |
| 3 | Sets up an isolated Python environment with a supported interpreter and current boto3 before any SageMaker deployment, training or AWS automation code runs. | huggingface/ | 11k | 2 repos | ~1.7k | Automated safety check: Pass | Apache-2.0 | 6 days ago |
| 4 | Operates the Inspire ML platform through its local `inspire` CLI: picking account, workspace and resources, launching notebooks, jobs and services, then cleaning up. | realZillionX/ | 548 | — | ~1.4k | Automated safety check: Pass | MIT | 10 days ago |
| 5 | Deploys SageMaker endpoints with autoscaling, CloudWatch alarms and tags on by default, using scripts for real-time, scale-to-zero and async setups. | huggingface/ | 11k | 1 repo | ~6.9k | Automated safety check: Pass | Apache-2.0 | 6 days ago |
| 6 | Designs ML pipeline infrastructure: experiment tracking with MLflow or Weights & Biases, Kubeflow and Airflow orchestration, Feast feature stores and model validation gates. | Jeffallan/ | 12k | 1 repo | ~1.9k | Automated safety check: Pass | MIT | 4 days ago |
| 7 | Guides an agent through tracking ML experiments with W&B: run logging, config capture, hyperparameter sweeps, artifacts and a model registry. | Orchestra-Research/ | 13k | 10 repos | ~3.1k | Automated safety check: Pass | MIT | 3 mo ago |
| 8 | Analyze a full code repository and generate one resume-ready project description grounded in repository evidence. | Ssabby1/ | 127 | — | ~1.8k | Automated safety check: Pass | MIT | 5 mo ago |
| 9 | Deploys open models or custom weights from Model Garden to Agent Platform endpoints, checks deployment status and cleans up endpoints, confirming before any change. | google/ | 21k | — | ~5.1k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 10 | Builds and maintains templates/index.mcp.json for Comfy Cloud MCP tools. | Comfy-Org/ | 1.3k | — | ~1.9k | Automated safety check: Notes | MIT | today |
| 11 | Runs ML training and batch jobs across clouds with SkyPilot, using spot instances, automatic region selection and managed recovery to cut GPU cost. | Orchestra-Research/ | 13k | 4 repos | ~2.4k | Automated safety check: Pass | MIT | 3 mo ago |
| 12 | Shows how to launch distributed Megatron-LM training on a SLURM cluster: sbatch skeleton, torch.distributed.run setup, CUDA_DEVICE_MAX_CONNECTIONS rules and failure diagnosis. | NVIDIA/ | 18k | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | today |
| 13 | Register a new AI model in DroidGear's model registry by fetching specs from models.dev. | Sunshow/ | 127 | — | ~2.5k | Automated safety check: Pass | MIT | today |
| 14 | Add, change or verify a Nexus LLM model definition — the registry entry, the provider default, and proof the model id actually works against the live endpoint. | ProfSynapse/ | 154 | — | ~1.1k | Automated safety check: Pass | MIT | 5 days ago |
| 15 | Declare the pipeline from data source to predictor as a skrub DataOps graph. | probabl-ai/ | 135 | — | ~4k | Automated safety check: Pass | BSD-3-Clause | yesterday |
| 16 | Sets up and maintains AzureML-ready Python projects as uv workspaces with devcontainers, a Makefile and job YAML, so local runs match cloud jobs and experiments stay reproducible. | Kilo-Org/ | 189 | — | ~3.1k | Automated safety check: Notes | MIT | 9 days ago |
| 17 | Guides an agent through designing an MLOps pipeline that covers data preparation, training, validation and deployment, with DAG orchestration and reference guides. | wshobson/ | 40k | 12 repos | ~1.8k | Automated safety check: Pass | MIT | 3 days ago |
| 18 | Covers logging and viewing training metrics, histograms, model graphs, embeddings and profiles with TensorBoard in PyTorch and TensorFlow projects. | Orchestra-Research/ | 13k | 3 repos | ~3.8k | Automated safety check: Pass | MIT | 3 mo ago |
| 19 | 19.Beava Design A skill your agent uses to generate well-branded interfaces and assets for beava (beava.dev, an open-source single-binary feature server for stream processing), either for production or throwaway… | beava-dev/ | 138 | — | ~443 | Automated safety check: Pass | Apache-2.0 | 4 mo ago |
| 20 | Evaluate one learner with skore.evaluate. An agent skill from probabl-ai/skills. | probabl-ai/ | 135 | — | ~4.2k | Automated safety check: Pass | BSD-3-Clause | yesterday |
| 21 | Guide to renting GPUs on Lambda Labs for ML training and inference: on-demand instances, 1-Click Clusters, SSH access, persistent filesystems and alternatives. | Orchestra-Research/ | 13k | 5 repos | ~3k | Automated safety check: Warn | MIT | 3 mo ago |
| 22 | Tracks ML experiments, versions models in the MLflow registry and covers deployment and reproducibility, with autologging for common frameworks. | Orchestra-Research/ | 13k | 2 repos | ~3.9k | Automated safety check: Pass | MIT | 3 mo ago |
| 23 | Shows how to log ML runs, configs, metrics and media with SwanLab and view them in cloud, local or self-hosted dashboards. | Orchestra-Research/ | 13k | 1 repo | ~2.4k | Automated safety check: Pass | MIT | 3 mo ago |
| 24 | 24.Edit how to use the edit command properly | omegaml/ | 107 | — | ~206 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 25 | Turns build, implement or design requests for ML pipelines into validated implementation plans grounded in a knowledge base or fetched framework documentation. | Leeroo-AI/ | 195 | — | ~11k | Automated safety check: Pass | Apache-2.0 | 6 mo ago |
| 26 | World-class ML engineering skill for productionizing ML models, MLOps, and building scalable ML systems. | davila7/ | 32k | 3 repos | ~1.4k | Automated safety check: Pass | MIT | today |
| 27 | Read-only audit of one persisted skore report: audit/NN<stem.py (jupytext percent), 1:1 with experiments/ and journal/. | probabl-ai/ | 135 | — | ~9.5k | Automated safety check: Pass | BSD-3-Clause | yesterday |
| 28 | Azure Machine Learning SDK v2 for Python. An agent skill from microsoft/skills. | microsoft/ | 3.1k | 6 repos | ~2.2k | Automated safety check: Pass | MIT | yesterday |
| 29 | ML engineering skill for productionizing models, building MLOps pipelines, and integrating LLMs. | alirezarezvani/ | 28k | 2 repos | ~2.4k | Automated safety check: Pass | MIT | 1 mo ago |
| 30 | Azure Weights & Biases SDK for .NET. An agent skill from microsoft/skills. | microsoft/ | 3.1k | 6 repos | ~2.8k | Automated safety check: Pass | MIT | yesterday |
| 31 | OCI Data Science service patterns including Jobs, Pipelines, Model Catalog, authentication, and the ADS SDK beyond AQUA. | oracle/ | 125 | — | ~2.1k | Automated safety check: Pass | UPL-1.0 | 1 mo ago |
| 32 | Data pipelines, feature stores, and embedding generation for AI/ML systems. | ancoleman/ | 526 | 1 repo | ~3.5k | Automated safety check: Pass | MIT | 10 mo ago |
| 33 | 33.ML Engineer Machine learning engineer expert for PyTorch, scikit-learn, model evaluation, and MLOps | RightNow-AI/ | 18k | — | ~987 | Automated safety check: Pass | Apache-2.0 | 3 mo ago |
| 34 | Build, test, and debug Hermes Agent RL environments for Atropos training. | Tommy-yw/ | 546 | — | ~3.3k | Automated safety check: Pass | MIT | 4 mo ago |
| 35 | Establish model registry standards, governance controls, metadata schemas, approvals, and lifecycle policies for enterprise AI deployments. | sickn33/ | 47k | 2 repos | ~3.9k | Automated safety check: Pass | MIT | yesterday |
| 36 | Strategic guidance for operationalizing machine learning models from experimentation to production. | ancoleman/ | 526 | 1 repo | ~9.2k | Automated safety check: Pass | MIT | 10 mo ago |
| 37 | Call model APIs through @prismshadow/mmsp (MMSP) — streaming text generation, image generation, speech synthesis, embeddings and the supported-model registry with one client. | Prism-Shadow/ | 2.5k | — | ~6.7k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 38 | Build comprehensive ML pipelines, experiment tracking, and model registries with MLflow, Kubeflow, and modern MLOps tools. | aiskillstore/ | 430 | 7 repos | ~2.8k | Automated safety check: Pass | No licence | today |
| 39 | 39.Nemo Curator Curate LLM training data: dedupe, filter, PII redaction. An agent skill from Luciole-Studio/Misaka-Agent. | Luciole-Studio/ | 125 | 2 repos | ~2.6k | Automated safety check: Pass | MIT | today |
| 40 | Agent Platform Model Registry Management. An agent skill from google/skills. | google/ | 21k | — | ~2.1k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 41 | Copy skore reports between local, Hub, and MLflow with skore sync, and optionally switch the recorded upload destination. | probabl-ai/ | 135 | — | ~1.6k | Automated safety check: Pass | BSD-3-Clause | yesterday |
| 42 | Use during planning, implementation, PR review, or /qv-qip-triage when a change may affect public SDK API, native dependency, plugin contract, model registry contract, runtime, transport, storage… | tetherto/ | 674 | — | ~746 | Automated safety check: Pass | Apache-2.0 | today |
| 43 | Neo4j Graph Data Science (GDS) embedded plugin via Python client or Cypher — covers graphdatascience client 2.x, GraphDataScience, gds.graph.project.native, gds.graph.project.cypher, snakecase… | neo4j-contrib/ | 114 | — | ~5k | Automated safety check: Notes | MIT | yesterday |
| 44 | Curated download URLs and target directories, organized by family (Flux, WAN, LTX, Qwen, Z-Image, SD15/SDXL), for every model the comfyui-mcp skills reference, covering checkpoints, VAEs, text… | artokun/ | 793 | — | ~2.1k | Automated safety check: Pass | MIT | 2 days ago |
| 45 | Design and implement a complete ML pipeline for: $ARGUMENTS. An agent skill from aiskillstore/marketplace. | aiskillstore/ | 430 | 7 repos | ~2.6k | Automated safety check: Pass | No licence | today |
| 46 | Shared review framework that every domain reviewer (pci, oracle, gov, edtech, healthcare, mlops, etc.) MUST follow. | avelikiy/ | 103 | — | ~3.1k | Automated safety check: Pass | MIT | yesterday |
| 47 | Read-only health check that a brand profile is production-ready: required fields, voice and audience completeness, guardrails, compliance-jurisdiction coverage, connector configuration and MCP… | indranilbanerjee/ | 854 | 1 repo | ~3.3k | Automated safety check: Notes | MIT | 3 days ago |
| 48 | Speed up long-sequence transformer training and inference. An agent skill from Luciole-Studio/Misaka-Agent. | Luciole-Studio/ | 125 | 1 repo | ~2.7k | Automated safety check: Pass | MIT | today |
Questions, answered from the data.
What is the best MLOps skill?
SageMaker IAM Role Preflight (official) from huggingface/skills ranks first of the 101 MLOps skills listed here, with the highest score: its repository has 11k GitHub stars, 1 other GitHub owner carry a copy, its SKILL.md loads about 1.8k tokens and it passes the automated safety check with no findings. Next come Py and Python Environment Setup for SageMaker.
Which MLOps skills are official?
11 of the 101 MLOps skills are official, published by the vendor's own GitHub organization: SageMaker IAM Role Preflight, Python Environment Setup for SageMaker, SageMaker Production Defaults, Model Garden Deployment, Megatron-LM on SLURM and 6 more.
How are these skills ranked?
By Skill Navigator score, which combines the GitHub stars of the skill's repository (shared across that repo's skills and discounted for large collections), how many other GitHub owners carry a copy of the skill, and automated SKILL.md quality checks, minus penalties for safety-check warnings and for each further skill from the same repository. Skills that fail the safety check are not listed.
Explore related skills
Category
More topics in DevOps & Cloud
- Deployment1,152
- CI/CD921
- Containers723
- Observability562
- Container orchestration519
- Infrastructure as code351
- Monitoring and alerting333
- Secrets management318
- Runbooks and postmortems277
- Incident response271
- Cloud networking230
- Backup and disaster recovery170
- Site reliability engineering150
- Cloud architecture114
- Cloud cost optimization93
- GitOps87
- Linux administration75
- Platform engineering51
- Chaos engineering28