Topic · DevOps & Cloud

Best MLOps skills for Claude Code, Codex and other agents.

Skills that train, deploy and monitor machine learning models in production.
skills
101
official
11

MLOps skills, ranked

Ranked by score. Sort bymost stars,trending,newest,recently updated

MLOps skills, ranked
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
1

Finds or validates a usable SageMaker execution role before deploying or training, so scripts do not try to create IAM roles they lack permission to create.

huggingface/skills11k1 repo~1.8kAutomated safety check: PassApache-2.06 days ago
2
2.Py

Comprehensive Python programming reference covering syntax, concurrency, networking, databases, ML/LLM development, and HPC.

crazyguitar/pysheeet8.2k—~886Automated safety check: PassMITyesterday
3

Sets up an isolated Python environment with a supported interpreter and current boto3 before any SageMaker deployment, training or AWS automation code runs.

huggingface/skills11k2 repos~1.7kAutomated safety check: PassApache-2.06 days ago
4

Operates the Inspire ML platform through its local `inspire` CLI: picking account, workspace and resources, launching notebooks, jobs and services, then cleaning up.

realZillionX/InspireSkill548—~1.4kAutomated safety check: PassMIT10 days ago
5

Deploys SageMaker endpoints with autoscaling, CloudWatch alarms and tags on by default, using scripts for real-time, scale-to-zero and async setups.

huggingface/skills11k1 repo~6.9kAutomated safety check: PassApache-2.06 days ago
6

Designs ML pipeline infrastructure: experiment tracking with MLflow or Weights & Biases, Kubeflow and Airflow orchestration, Feast feature stores and model validation gates.

Jeffallan/claude-skills12k1 repo~1.9kAutomated safety check: PassMIT4 days ago
7

Guides an agent through tracking ML experiments with W&B: run logging, config capture, hyperparameter sweeps, artifacts and a model registry.

Orchestra-Research/AI-Research-SKILLs13k10 repos~3.1kAutomated safety check: PassMIT3 mo ago
8

Analyze a full code repository and generate one resume-ready project description grounded in repository evidence.

Ssabby1/repo-to-resume-tailor127—~1.8kAutomated safety check: PassMIT5 mo ago
9

Deploys open models or custom weights from Model Garden to Agent Platform endpoints, checks deployment status and cleans up endpoints, confirming before any change.

google/skills21k—~5.1kAutomated safety check: PassApache-2.0yesterday
10

Builds and maintains templates/index.mcp.json for Comfy Cloud MCP tools.

Comfy-Org/workflow_templates1.3k—~1.9kAutomated safety check: NotesMITtoday
11

Runs ML training and batch jobs across clouds with SkyPilot, using spot instances, automatic region selection and managed recovery to cut GPU cost.

Orchestra-Research/AI-Research-SKILLs13k4 repos~2.4kAutomated safety check: PassMIT3 mo ago
12

Shows how to launch distributed Megatron-LM training on a SLURM cluster: sbatch skeleton, torch.distributed.run setup, CUDA_DEVICE_MAX_CONNECTIONS rules and failure diagnosis.

NVIDIA/Megatron-LM18k—~1.8kAutomated safety check: PassApache-2.0today
13

Register a new AI model in DroidGear's model registry by fetching specs from models.dev.

Sunshow/droidgear127—~2.5kAutomated safety check: PassMITtoday
14

Add, change or verify a Nexus LLM model definition — the registry entry, the provider default, and proof the model id actually works against the live endpoint.

ProfSynapse/nexus154—~1.1kAutomated safety check: PassMIT5 days ago
15

Declare the pipeline from data source to predictor as a skrub DataOps graph.

probabl-ai/skills135—~4kAutomated safety check: PassBSD-3-Clauseyesterday
16

Sets up and maintains AzureML-ready Python projects as uv workspaces with devcontainers, a Makefile and job YAML, so local runs match cloud jobs and experiments stay reproducible.

Kilo-Org/kilo-marketplace189—~3.1kAutomated safety check: NotesMIT9 days ago
17

Guides an agent through designing an MLOps pipeline that covers data preparation, training, validation and deployment, with DAG orchestration and reference guides.

wshobson/agents40k12 repos~1.8kAutomated safety check: PassMIT3 days ago
18

Covers logging and viewing training metrics, histograms, model graphs, embeddings and profiles with TensorBoard in PyTorch and TensorFlow projects.

Orchestra-Research/AI-Research-SKILLs13k3 repos~3.8kAutomated safety check: PassMIT3 mo ago
19

A skill your agent uses to generate well-branded interfaces and assets for beava (beava.dev, an open-source single-binary feature server for stream processing), either for production or throwaway…

beava-dev/beava138—~443Automated safety check: PassApache-2.04 mo ago
20

Evaluate one learner with skore.evaluate. An agent skill from probabl-ai/skills.

probabl-ai/skills135—~4.2kAutomated safety check: PassBSD-3-Clauseyesterday
21

Guide to renting GPUs on Lambda Labs for ML training and inference: on-demand instances, 1-Click Clusters, SSH access, persistent filesystems and alternatives.

Orchestra-Research/AI-Research-SKILLs13k5 repos~3kAutomated safety check: WarnMIT3 mo ago
22

Tracks ML experiments, versions models in the MLflow registry and covers deployment and reproducibility, with autologging for common frameworks.

Orchestra-Research/AI-Research-SKILLs13k2 repos~3.9kAutomated safety check: PassMIT3 mo ago
23

Shows how to log ML runs, configs, metrics and media with SwanLab and view them in cloud, local or self-hosted dashboards.

Orchestra-Research/AI-Research-SKILLs13k1 repo~2.4kAutomated safety check: PassMIT3 mo ago
24
24.Edit

how to use the edit command properly

omegaml/omegaml107—~206Automated safety check: PassApache-2.0yesterday
25

Turns build, implement or design requests for ML pipelines into validated implementation plans grounded in a knowledge base or fetched framework documentation.

Leeroo-AI/superml195—~11kAutomated safety check: PassApache-2.06 mo ago
26

World-class ML engineering skill for productionizing ML models, MLOps, and building scalable ML systems.

davila7/claude-code-templates32k3 repos~1.4kAutomated safety check: PassMITtoday
27

Read-only audit of one persisted skore report: audit/NN<stem.py (jupytext percent), 1:1 with experiments/ and journal/.

probabl-ai/skills135—~9.5kAutomated safety check: PassBSD-3-Clauseyesterday
28
28.Azure AI ML PyOfficial

Azure Machine Learning SDK v2 for Python. An agent skill from microsoft/skills.

microsoft/skills3.1k6 repos~2.2kAutomated safety check: PassMITyesterday
29

ML engineering skill for productionizing models, building MLOps pipelines, and integrating LLMs.

alirezarezvani/claude-skills28k2 repos~2.4kAutomated safety check: PassMIT1 mo ago
30

Azure Weights & Biases SDK for .NET. An agent skill from microsoft/skills.

microsoft/skills3.1k6 repos~2.8kAutomated safety check: PassMITyesterday
31
31.Oci Data ScienceOfficial

OCI Data Science service patterns including Jobs, Pipelines, Model Catalog, authentication, and the ADS SDK beyond AQUA.

oracle/accelerated-data-science125—~2.1kAutomated safety check: PassUPL-1.01 mo ago
32

Data pipelines, feature stores, and embedding generation for AI/ML systems.

ancoleman/ai-design-components5261 repo~3.5kAutomated safety check: PassMIT10 mo ago
33

Machine learning engineer expert for PyTorch, scikit-learn, model evaluation, and MLOps

RightNow-AI/openfang18k—~987Automated safety check: PassApache-2.03 mo ago
34

Build, test, and debug Hermes Agent RL environments for Atropos training.

Tommy-yw/RunbookHermes546—~3.3kAutomated safety check: PassMIT4 mo ago
35

Establish model registry standards, governance controls, metadata schemas, approvals, and lifecycle policies for enterprise AI deployments.

sickn33/agentic-awesome-skills47k2 repos~3.9kAutomated safety check: PassMITyesterday
36

Strategic guidance for operationalizing machine learning models from experimentation to production.

ancoleman/ai-design-components5261 repo~9.2kAutomated safety check: PassMIT10 mo ago
37

Call model APIs through @prismshadow/mmsp (MMSP) — streaming text generation, image generation, speech synthesis, embeddings and the supported-model registry with one client.

Prism-Shadow/penguin-harness2.5k—~6.7kAutomated safety check: PassApache-2.0yesterday
38

Build comprehensive ML pipelines, experiment tracking, and model registries with MLflow, Kubeflow, and modern MLOps tools.

aiskillstore/marketplace4307 repos~2.8kAutomated safety check: PassNo licencetoday
39

Curate LLM training data: dedupe, filter, PII redaction. An agent skill from Luciole-Studio/Misaka-Agent.

Luciole-Studio/Misaka-Agent1252 repos~2.6kAutomated safety check: PassMITtoday
40

Agent Platform Model Registry Management. An agent skill from google/skills.

google/skills21k—~2.1kAutomated safety check: PassApache-2.0yesterday
41

Copy skore reports between local, Hub, and MLflow with skore sync, and optionally switch the recorded upload destination.

probabl-ai/skills135—~1.6kAutomated safety check: PassBSD-3-Clauseyesterday
42

Use during planning, implementation, PR review, or /qv-qip-triage when a change may affect public SDK API, native dependency, plugin contract, model registry contract, runtime, transport, storage…

tetherto/qvac674—~746Automated safety check: PassApache-2.0today
43

Neo4j Graph Data Science (GDS) embedded plugin via Python client or Cypher — covers graphdatascience client 2.x, GraphDataScience, gds.graph.project.native, gds.graph.project.cypher, snakecase…

neo4j-contrib/neo4j-skills114—~5kAutomated safety check: NotesMITyesterday
44

Curated download URLs and target directories, organized by family (Flux, WAN, LTX, Qwen, Z-Image, SD15/SDXL), for every model the comfyui-mcp skills reference, covering checkpoints, VAEs, text…

artokun/comfyui-mcp793—~2.1kAutomated safety check: PassMIT2 days ago
45

Design and implement a complete ML pipeline for: $ARGUMENTS. An agent skill from aiskillstore/marketplace.

aiskillstore/marketplace4307 repos~2.6kAutomated safety check: PassNo licencetoday
46

Shared review framework that every domain reviewer (pci, oracle, gov, edtech, healthcare, mlops, etc.) MUST follow.

avelikiy/great_cto103—~3.1kAutomated safety check: PassMITyesterday
47

Read-only health check that a brand profile is production-ready: required fields, voice and audience completeness, guardrails, compliance-jurisdiction coverage, connector configuration and MCP…

indranilbanerjee/digital-marketing-pro8541 repo~3.3kAutomated safety check: NotesMIT3 days ago
48

Speed up long-sequence transformer training and inference. An agent skill from Luciole-Studio/Misaka-Agent.

Luciole-Studio/Misaka-Agent1251 repo~2.7kAutomated safety check: PassMITtoday

Questions, answered from the data.

What is the best MLOps skill?

SageMaker IAM Role Preflight (official) from huggingface/skills ranks first of the 101 MLOps skills listed here, with the highest score: its repository has 11k GitHub stars, 1 other GitHub owner carry a copy, its SKILL.md loads about 1.8k tokens and it passes the automated safety check with no findings. Next come Py and Python Environment Setup for SageMaker.

Which MLOps skills are official?

11 of the 101 MLOps skills are official, published by the vendor's own GitHub organization: SageMaker IAM Role Preflight, Python Environment Setup for SageMaker, SageMaker Production Defaults, Model Garden Deployment, Megatron-LM on SLURM and 6 more.

How are these skills ranked?

By Skill Navigator score, which combines the GitHub stars of the skill's repository (shared across that repo's skills and discounted for large collections), how many other GitHub owners carry a copy of the skill, and automated SKILL.md quality checks, minus penalties for safety-check warnings and for each further skill from the same repository. Skills that fail the safety check are not listed.