Topic · DevOps & Cloud
Best MLOps skills, page 2
MLOps skills, ranked
Ranked by score. Sort bymost stars,trending,newest,recently updated
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 49 | 49.Saelens Train sparse autoencoders to interpret model features. An agent skill from Luciole-Studio/Misaka-Agent. | Luciole-Studio/ | 139 | 1 repo | ~3.7k | Automated safety check: Pass | MIT | today |
| 50 | 50.Simpo Reference-free preference alignment, simpler than DPO. An agent skill from Luciole-Studio/Misaka-Agent. | Luciole-Studio/ | 139 | 1 repo | ~1.4k | Automated safety check: Pass | MIT | today |
| 51 | 51.Slime RL post-training for LLMs with Megatron and SGLang. An agent skill from Luciole-Studio/Misaka-Agent. | Luciole-Studio/ | 139 | 1 repo | ~2.7k | Automated safety check: Pass | MIT | today |
| 52 | 52.Tensorrt LLM High-throughput LLM inference on NVIDIA GPUs. An agent skill from Luciole-Studio/Misaka-Agent. | Luciole-Studio/ | 139 | 1 repo | ~1.3k | Automated safety check: Pass | MIT | today |
| 53 | 53.Torchtitan Pretrain LLMs at scale with PyTorch 4D parallelism. An agent skill from Luciole-Studio/Misaka-Agent. | Luciole-Studio/ | 139 | 1 repo | ~2.6k | Automated safety check: Pass | MIT | today |
| 54 | Read-only health check that a brand profile is production-ready: required fields, voice and audience completeness, guardrails, compliance-jurisdiction coverage, connector configuration and MCP… | indranilbanerjee/ | 855 | 1 repo | ~3.3k | Automated safety check: Notes | MIT | 4 days ago |
| 55 | 55.AI ML AI and machine learning workflow covering LLM application development, RAG implementation, agent architecture, ML pipelines, and AI-powered features. | aiskillstore/ | 430 | 4 repos | ~1.5k | Automated safety check: Pass | No licence | today |
| 56 | Train ML models on Databricks. An agent skill from databricks/databricks-agent-skills. | databricks/ | 345 | — | ~4.6k | Automated safety check: Pass | Unknown | today |
| 57 | Execute feature store connector operations. An agent skill from jeremylongshore/tons-of-skills-marketplace. | jeremylongshore/ | 2.8k | 1 repo | ~574 | Automated safety check: Pass | MIT | today |
| 58 | Manage model registry manager operations. An agent skill from jeremylongshore/tons-of-skills-marketplace. | jeremylongshore/ | 2.8k | 1 repo | ~569 | Automated safety check: Pass | MIT | today |
| 59 | Automate ML workflows with Airflow, Kubeflow, MLflow. An agent skill from secondsky/claude-skills. | secondsky/ | 227 | 1 repo | ~3.2k | Automated safety check: Pass | MIT | 9 days ago |
| 60 | Deploy production recommendation systems with feature stores, caching, A/B testing. | secondsky/ | 227 | 1 repo | ~3.6k | Automated safety check: Pass | MIT | 9 days ago |
| 61 | Neo4j Graph Data Science (GDS) embedded plugin via Python client or Cypher — covers graphdatascience client 2.x, GraphDataScience, gds.graph.project.native, gds.graph.project.cypher, snakecase… | neo4j-contrib/ | 114 | — | ~5k | Automated safety check: Notes | MIT | yesterday |
| 62 | Execute use when provisioning Vertex AI infrastructure with Terraform. | jeremylongshore/ | 2.8k | — | ~709 | Automated safety check: Pass | MIT | today |
| 63 | Shared review framework that every domain reviewer (pci, oracle, gov, edtech, healthcare, mlops, etc.) MUST follow. | avelikiy/ | 103 | — | ~3.1k | Automated safety check: Pass | MIT | today |
| 64 | A Principal ML Engineer interviewer that simulates a FAANG-style ML system design interview covering the full lifecycle from data to production. | PrepLabsAI/ | 112 | — | ~4.2k | Automated safety check: Pass | MIT | yesterday |
| 65 | Re-create the local 3-hour tri-cluster cluster-sweep loop (Leonardo + CoreWeave(iris) + TACC(Vista); Jupiter SKIPPED until ~Jul 12) — the autonomous ML-ops monitor — if it has been lost. | open-thoughts/ | 301 | — | ~4.1k | Automated safety check: Pass | Apache-2.0 | 9 days ago |
| 66 | Pick how to write a figure before custom plot code. An agent skill from probabl-ai/skills. | probabl-ai/ | 137 | — | ~785 | Automated safety check: Pass | BSD-3-Clause | today |
| 67 | Literature and web research for an ML methodology concern (EDA extra measurements, leakage, transforms, feature engineering, learner family), or an EDA extra-analysis survey from JOURNAL plus… | probabl-ai/ | 137 | — | ~1.1k | Automated safety check: Pass | BSD-3-Clause | today |
| 68 | Selects, deploys, and customizes AI models on Amazon SageMaker. | aws/ | 2.8k | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | today |
| 69 | MLOps across model deployment, ML pipelines, monitoring, and feature stores. | borghei/ | 881 | — | ~3.5k | Automated safety check: Pass | MIT | today |
| 70 | 70.Lambda Labs On-demand GPU cloud instances for ML training. An agent skill from Luciole-Studio/Misaka-Agent. | Luciole-Studio/ | 139 | 2 repos | ~3k | Automated safety check: Warn | MIT | today |
| 71 | Store reviews and in-app updates on the KMP Starter Template — StarterStoreManager (askForReview, checkAppUpdate), rememberStarterStoreManager, and AppUpdateProvider. | DevAtrii/ | 167 | — | ~670 | Automated safety check: Pass | MIT | 1 mo ago |
| 72 | Owns one experiment's smoke test: a small pytest that fits on part of real data/ and predicts on a disjoint slice with no pre-history buffer. | probabl-ai/ | 137 | — | ~6k | Automated safety check: Pass | BSD-3-Clause | today |
| 73 | Guided workflow for taking any ML codebase — including one with no experiment tracking at all, or one full of MLflow 2-era idioms — to a production-grade open-source MLflow 3 setup with… | pproenca/ | 215 | — | ~1.9k | Automated safety check: Pass | MIT | 1 mo ago |
| 74 | Quantify prediction uncertainty of MACE MLIPs using committee (ensemble) models; flag high-uncertainty structures for DFT verification. | learningmatter-mit/ | 176 | — | ~2.3k | Automated safety check: Pass | MIT | today |
| 75 | A skill your agent uses when the user asks to "design an experiment", "build a predictive model", "run A/B test analysis", "perform causal inference", "engineer features", "evaluate model… | borghei/ | 881 | — | ~1.7k | Automated safety check: Pass | MIT | today |
| 76 | Canonical backlog loop step. An agent skill from probabl-ai/skills. | probabl-ai/ | 137 | — | ~3.3k | Automated safety check: Pass | BSD-3-Clause | today |
| 77 | 77.ML Ops A skill your agent uses when deploying ML models to production, setting up model monitoring, implementing A/B testing for models, or managing feature stores. | majiayu000/ | 666 | 1 repo | ~4.3k | Automated safety check: Pass | MIT | today |
| 78 | Design, implement, and validate reproducible machine-learning pipelines spanning data preparation, training, evaluation, registry, and deployment gates. | majiayu000/ | 666 | 1 repo | ~1.5k | Automated safety check: Pass | MIT | today |
| 79 | Expert in Machine Learning Operations bridging data science and DevOps. | majiayu000/ | 666 | 1 repo | ~820 | Automated safety check: Pass | MIT | today |
| 80 | 80.Pyhealth Build clinical/healthcare deep-learning pipelines with PyHealth — loading EHR/signal/imaging datasets (MIMIC-III/IV, eICU, OMOP, SleepEDF, ChestXray14, EHRShot), defining tasks (mortality… | majiayu000/ | 666 | 1 repo | ~1.8k | Automated safety check: Pass | MIT | today |
| 81 | Google Cloud Vertex AI for enterprise Gemini deployments — production scaling, fine-tuning, and MLOps. | majiayu000/ | 666 | 1 repo | ~2.1k | Automated safety check: Pass | Apache-2.0 | today |
| 82 | 82.AI ML V2 AI/ML Workflow Bundle workflow skill. An agent skill from majiayu000/claude-skill-registry. | majiayu000/ | 666 | 1 repo | ~3.4k | Automated safety check: Pass | MIT | today |
| 83 | World-Class Technology & Data Playbook. An agent skill from LeoYeAI/openclaw-master-skills. | LeoYeAI/ | 2.2k | — | ~6.8k | Automated safety check: Pass | MIT | 2 mo ago |
| 84 | Guide for experiment tracking tool setup (MLflow, Weights & Biases, etc.), reproducibility assurance, model registry, and experiment comparison methodology. | revfactory/ | 1.3k | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | 6 mo ago |
| 85 | 85.Cortex Model Build an ML pipeline — from data to trained model to serving endpoint. | jeremylongshore/ | 2.8k | — | ~1.2k | Automated safety check: Notes | MIT | today |
| 86 | A skill your agent uses whenever the user wants to find, shortlist, vet, or enrich US AI/ML/data consulting firms (consultancies) — AI/ML development, MLOps, generative AI / LLM apps (RAG, chatbots… | jeremylongshore/ | 2.8k | — | ~4k | Automated safety check: Notes | MIT | today |
| 87 | A skill your agent uses whenever the user wants to find, shortlist, vet, or enrich US software development firms — custom software, web development, mobile app development, backend/API development… | jeremylongshore/ | 2.8k | — | ~3.7k | Automated safety check: Notes | MIT | today |
| 88 | Papers on LLMs for IT operations and AIOps research. An agent skill from wentorai/research-plugins. | wentorai/ | 298 | 1 repo | ~3.3k | Automated safety check: Pass | MIT | 3 mo ago |
| 89 | Build and deploy reproducible production ML pipelines for research | wentorai/ | 298 | 1 repo | ~2.3k | Automated safety check: Pass | MIT | 3 mo ago |
| 90 | Coaches end-to-end ML system design interviews covering inference pipelines, recommendation systems, RAG, feature stores, and monitoring. | curiositech/ | 243 | — | ~3.4k | Automated safety check: Pass | MIT | 1 mo ago |
| 91 | Deploy ML models with FastAPI, Docker, Kubernetes. An agent skill from secondsky/claude-skills. | secondsky/ | 227 | — | ~2.4k | Automated safety check: Pass | MIT | 9 days ago |
| 92 | 92.Feast A skill your agent uses for Feast feature store tasks: feature repositories, definitions, CLI, retrieval, materialization, serving, RAG/vector search, integrations, and Feast contributor workflows. | VectorSpaceLab/ | 328 | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 93 | Guide for selecting the most appropriate foundation MLIP model based on simulation requirements. | learningmatter-mit/ | 176 | — | ~1.5k | Automated safety check: Pass | MIT | today |
| 94 | GPU-accelerate Python code using CuPy, Numba CUDA, Warp, cuDF, cuML, cuGraph, KvikIO, cuCIM, cuxfilter, cuVS, cuSpatial, and RAFT. | majiayu000/ | 666 | 1 repo | ~8.5k | Automated safety check: Pass | MIT | today |
| 95 | Guide for maintaining the MassGen model and backend registry. | Microck/ | 403 | — | ~3.6k | Automated safety check: Pass | Unknown | 1 mo ago |
| 96 | ML engineering skill for productionizing models, building MLOps pipelines, and integrating LLMs. | borghei/ | 881 | — | ~1.7k | Automated safety check: Pass | MIT | today |
Explore related skills
Category
More topics in DevOps & Cloud
- Deployment1,264
- CI/CD977
- Containers731
- Observability636
- Container orchestration542
- Infrastructure as code377
- Monitoring and alerting351
- Secrets management335
- Runbooks and postmortems324
- Incident response302
- Cloud networking225
- Backup and disaster recovery183
- Site reliability engineering152
- Cloud architecture118
- Cloud cost optimization92
- GitOps89
- Linux administration70
- Platform engineering45
- Chaos engineering25