Search
DevOps & Cloud · For data scientists
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones. | huggingface/ | 11k | 1 repo | ~4.6k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 2 | Decide, don't guess — trigger on ANY combinatorial or ground-state decision where a plausible guess is worse than silence: rosters and on-call schedules, packing and placement, RAG passage… | brayonpi/ | 1.4k | — | ~5k | Automated safety check: Pass | Proprietary | 1 mo ago |
| 3 | Operates the Inspire ML platform through its local `inspire` CLI: picking account, workspace and resources, launching notebooks, jobs and services, then cleaning up. | realZillionX/ | 551 | — | ~1.4k | Automated safety check: Pass | MIT | 2 days ago |
| 4 | Integrate a benchmark or custom environment into SAfactory using fixed adapter templates and local contract tests, optionally run Docker/RJob evaluation, or prepare GRPO/RL training. | AI45Lab/ | 236 | — | ~1.8k | Automated safety check: Pass | No licence | 16 days ago |
| 5 | 5.Langfuse Interact with Langfuse and access its documentation: tracing, monitoring, creating datasets, running experiments, and evaluating AI applications. | langfuse/ | 301 | — | ~2.1k | Automated safety check: Notes | MIT | 9 days ago |
| 6 | Diagnose failed, hung, slow, OOM, NaN, illegal-memory-access, NCCL, compilation, rollout, or training runs in AReno. | inclusionAI/ | 323 | — | ~486 | Automated safety check: Pass | Apache-2.0 | today |
| 7 | Shows how to launch distributed Megatron-LM training on a SLURM cluster: sbatch skeleton, torch.distributed.run setup, CUDA_DEVICE_MAX_CONNECTIONS rules and failure diagnosis. | NVIDIA/ | 18k | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | today |
| 8 | A skill your agent uses when writing, editing, reviewing, or running Helm chart tests for the Astronomer airflow-chart repository. | astronomer/ | 297 | — | ~2.8k | Automated safety check: Pass | Unknown | 2 days ago |
| 9 | Plans memory headroom, works through out-of-memory failures and watches temperature and power during long ML training jobs on NVIDIA DGX Spark. | wshobson/ | 40k | — | ~2k | Automated safety check: Pass | MIT | 5 days ago |
| 10 | 10.Eval Evaluate and score agent behavior against a golden reference. | agentevals-dev/ | 163 | — | ~904 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 11 | Build an auditable release-evidence workflow for a desktop or packaged application. | Ali-Marandi/ | 107 | — | ~1.6k | Automated safety check: Pass | MIT | 1 mo ago |
| 12 | Guides an agent through tracking ML experiments with W&B: run logging, config capture, hyperparameter sweeps, artifacts and a model registry. | Orchestra-Research/ | 13k | 9 repos | ~3.1k | Automated safety check: Pass | MIT | 3 mo ago |
| 13 | Submit, monitor and benchmark jax-gcm (jcm) simulations on NCAR Derecho's PBS queues. | climate-analytics-lab/ | 108 | — | ~2.5k | Automated safety check: Pass | Apache-2.0 | today |
| 14 | Generate preventive Well-Architected guardrails — AWS Config rules, Service Control Policies, permission boundaries, CloudWatch alarms, and IaC policy checks (CDK Aspects, cfn-guard, OPA/Sentinel) —… | aws-samples/ | 275 | — | ~2.8k | Automated safety check: Pass | MIT-0 | 4 days ago |
| 15 | 15.Budget Set Define a spend budget for Claude Code and, optionally, create a cost alert rule that fires when usage crosses the limit, via POST /api/alerts/rules on the Agent Monitor dashboard. | hoangsonww/ | 1.1k | — | ~1k | Automated safety check: Pass | MIT | today |
| 16 | Logs and visualizes ML training metrics with Trackio, firing alerts for issues like loss spikes, and syncing a live dashboard to a Hugging Face Space. | huggingface/ | 11k | 2 repos | ~1.3k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 17 | Runs ML training and batch jobs across clouds with SkyPilot, using spot instances, automatic region selection and managed recovery to cut GPU cost. | Orchestra-Research/ | 13k | 4 repos | ~2.4k | Automated safety check: Pass | MIT | 3 mo ago |
| 18 | A skill your agent uses when writing, editing, reviewing, or running functional (end-to-end) tests for the Astronomer airflow-chart repository. | astronomer/ | 297 | — | ~2.2k | Automated safety check: Pass | Unknown | 2 days ago |
| 19 | 19.Vertex AI Primary Router for Vertex AI skills. An agent skill from GoogleCloudPlatform/vertex-ai-samples. | GoogleCloudPlatform/ | 792 | — | ~522 | Automated safety check: Pass | Apache-2.0 | today |
| 20 | Create a new Agent Skill following project standards and templates. | oocx/ | 174 | — | ~1.5k | Automated safety check: Pass | MIT | today |
| 21 | Deploys Airflow DAGs and projects. An agent skill from astronomer/agents. | astronomer/ | 451 | 1 repo | ~2.8k | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 22 | Sets up and maintains AzureML-ready Python projects as uv workspaces with devcontainers, a Makefile and job YAML, so local runs match cloud jobs and experiments stay reproducible. | Kilo-Org/ | 190 | — | ~3.1k | Automated safety check: Notes | MIT | 11 days ago |
| 23 | Guides an agent through designing an MLOps pipeline that covers data preparation, training, validation and deployment, with DAG orchestration and reference guides. | wshobson/ | 40k | 12 repos | ~1.8k | Automated safety check: Pass | MIT | 5 days ago |
| 24 | OpenVINO — real-time object detection via Docker (NCS2, Intel GPU, CPU) | SharpAI/ | 3.1k | — | ~1.3k | Automated safety check: Pass | MIT | 23 days ago |
| 25 | Autonomous AI agent platform for building and deploying continuous agents. | Orchestra-Research/ | 13k | 2 repos | ~2.3k | Automated safety check: Notes | MIT | 3 mo ago |
| 26 | LLM observability platform for tracing, evaluation, and monitoring. | Orchestra-Research/ | 13k | 2 repos | ~2.4k | Automated safety check: Pass | MIT | 3 mo ago |
| 27 | Sets up Arize Phoenix to trace, evaluate and monitor LLM applications, with instrumentation for OpenAI, LangChain and LlamaIndex and a self-hosted server. | Orchestra-Research/ | 13k | 2 repos | ~2.9k | Automated safety check: Pass | MIT | 3 mo ago |
| 28 | 28.Deploy Deploy functions, hosting, or full release to Firebase. An agent skill from duyet/pricetrack. | duyet/ | 139 | — | ~338 | Automated safety check: Pass | MIT | 5 days ago |
| 29 | 29.Helm Chart A skill your agent uses for Helm chart work - creating charts, modifying existing charts, values design, testing. | astronomer/ | 297 | — | ~6.4k | Automated safety check: Pass | Unknown | 2 days ago |
| 30 | Tracks ML experiments, versions models in the MLflow registry and covers deployment and reproducibility, with autologging for common frameworks. | Orchestra-Research/ | 13k | 2 repos | ~3.9k | Automated safety check: Pass | MIT | 3 mo ago |
| 31 | When the user reports a bug, error message, stack trace, unexpected behavior, or "why is this broken" question, search the corpus first for prior occurrences, known fixes, or related runbooks. | lyonzin/ | 292 | — | ~1.8k | Automated safety check: Pass | MIT | yesterday |
| 32 | Entry point for hosting a model on Amazon SageMaker: asks a few questions, picks a deployment pathway and hands off to the specialist skills. | huggingface/ | 11k | 1 repo | ~2.1k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 33 | Analyzes BigQuery slot use, query costs and execution bottlenecks from INFORMATION_SCHEMA to diagnose slow queries, slot contention and unpartitioned scans. | google/ | 21k | — | ~2.3k | Automated safety check: Pass | Apache-2.0 | today |
| 34 | Guide to renting GPUs on Lambda Labs for ML training and inference: on-demand instances, 1-Click Clusters, SSH access, persistent filesystems and alternatives. | Orchestra-Research/ | 13k | 4 repos | ~3k | Automated safety check: Warn | MIT | 3 mo ago |
| 35 | INVOKE THIS SKILL when adding Arize AX tracing or observability to an app for the first time, or when the user wants to instrument their LLM app or get started with LLM observability. | boshi-xixixi/ | 276 | — | ~5.1k | Automated safety check: Notes | MIT | 5 mo ago |
| 36 | Decide whether and how errors report to Sentry. An agent skill from langfuse/langfuse. | langfuse/ | 36k | — | ~3k | Automated safety check: Pass | Unknown | today |
| 37 | Run or install repo security leak checks with BetterLeaks and Trivy. | instructa/ | 139 | — | ~557 | Automated safety check: Pass | No licence | 12 days ago |
| 38 | Shared workflow for editing Langfuse's repo-owned agent setup under .agents/. | langfuse/ | 36k | — | ~799 | Automated safety check: Pass | Unknown | today |
| 39 | Designs ML pipeline infrastructure: experiment tracking with MLflow or Weights & Biases, Kubeflow and Airflow orchestration, Feast feature stores and model validation gates. | Jeffallan/ | 12k | — | ~1.9k | Automated safety check: Pass | MIT | 7 days ago |
| 40 | World-class ML engineering skill for productionizing ML models, MLOps, and building scalable ML systems. | davila7/ | 33k | 2 repos | ~1.4k | Automated safety check: Pass | MIT | today |
| 41 | A skill your agent uses to deploy the vss-behavior-analytics service standalone (entrypoint, config-source, optional calibration). | NVIDIA-AI-Blueprints/ | 1.9k | — | ~2.7k | Automated safety check: Pass | Apache-2.0 | today |
| 42 | Expert in building products that wrap AI APIs (OpenAI, Anthropic, etc.) into focused tools people will pay for. | davila7/ | 33k | 4 repos | ~1.7k | Automated safety check: Pass | MIT | today |
| 43 | WORKFLOW SKILL — Creates, restructures, and audits GitHub Copilot .agent.md and .prompt.md files with correct frontmatter, handoffs, model policy, context budgets, and validation. | jonathan-vella/ | 217 | — | ~1.8k | Automated safety check: Pass | MIT | today |
| 44 | Azure Maps SDK for .NET. An agent skill from microsoft/skills. | microsoft/ | 3.1k | 5 repos | ~3.9k | Automated safety check: Pass | MIT | yesterday |
| 45 | OCI Data Science service patterns including Jobs, Pipelines, Model Catalog, authentication, and the ADS SDK beyond AQUA. | oracle/ | 125 | — | ~2.1k | Automated safety check: Pass | UPL-1.0 | 1 mo ago |
| 46 | Answers "for this model, this hardware, this GPU budget and this workload, what is the best known serving configuration, and how much do I trust it?" by walking an ordered set of recipe catalogs… | ai-dynamo/ | 8.3k | — | ~4.3k | Automated safety check: Pass | Apache-2.0 | today |
| 47 | 47.Mle Workflow Production machine-learning engineering workflow for data contracts, reproducible training, model evaluation, deployment, monitoring, and rollback. | affaan-m/ | 276k | 1 repo | ~5.6k | Automated safety check: Pass | MIT | today |
| 48 | Select an on-device OpenMed PII model from the committed registry by language, runtime format, and size budget, then require recall validation before deployment. | maziyarpanahi/ | 5.5k | — | ~798 | Automated safety check: Pass | Apache-2.0 | today |