Search

DevOps & Cloud · vLLM

26 skills found.
Search results
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
1

Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.

huggingface/skills11k1 repo~4.6kAutomated safety check: PassApache-2.02 days ago
2

Runs the end-to-end vLLM Ascend release process: opens the release checklist and feedback issues, scans for release-blocking bugs and test coverage gaps, and generates release notes and announcements.

vllm-project/vllm-ascend2.9k—~7.2kAutomated safety check: PassApache-2.0today
3

Deploys SageMaker endpoints with autoscaling, CloudWatch alarms and tags on by default, using scripts for real-time, scale-to-zero and async setups.

huggingface/skills11k1 repo~6.9kAutomated safety check: PassApache-2.02 days ago
4

Plans memory headroom, works through out-of-memory failures and watches temperature and power during long ML training jobs on NVIDIA DGX Spark.

wshobson/agents40k—~2kAutomated safety check: PassMIT5 days ago
5

Deploys LLMs with vLLM for high-throughput serving, covering the OpenAI-compatible server, offline batch inference, monitoring and a Docker rollout.

Orchestra-Research/AI-Research-SKILLs13k5 repos~2.3kAutomated safety check: PassMIT3 mo ago
6

Advisory guidance for Amazon EKS architecture and configuration decisions — compute strategy, networking, security, reliability, cost, autoscaling, observability, multi-tenancy, and upgrade planning.

aws-samples/appmod-blueprints115—~5kAutomated safety check: PassMIT-02 days ago
7

Compose, edit, refactor, and validate Lego-RL train/eval/infer .env configs and reusable scripts/templates modules.

LegoX/Lego-RL113—~2.1kAutomated safety check: NotesApache-2.02 days ago
8

Interactively build, push or load, and deploy an airunway component (controller or any provider) to the cluster

ai-runway/airunway102—~927Automated safety check: PassApache-2.013 days ago
9

Deploy vLLM to Kubernetes (K8s) with GPU support, health probes, and OpenAI-compatible API endpoint.

vllm-project/vllm-skills102—~2kAutomated safety check: PassApache-2.06 mo ago
10

Deploys and manages VSS through setup.sh and its Docker Compose overlays.

open-edge-platform/edge-ai-libraries171—~4.1kAutomated safety check: PassApache-2.0today
11

Upgrade focused runtime dependencies in AReaL. An agent skill from areal-project/AReaL.

areal-project/AReaL5.8k—~6kAutomated safety check: PassApache-2.0today
12

A skill your agent uses whenever a developer needs to deploy VSS to Kubernetes, helm install VSS, configure values.yaml for VSS, or run VSS on k8s with GPU/vLLM for the…

open-edge-platform/edge-ai-libraries171—~3.8kAutomated safety check: PassApache-2.0today
13
13.Aqua MetricsOfficial

Set up Prometheus and Grafana monitoring for AQUA vLLM model deployments on OCI.

oracle/accelerated-data-science125—~1.5kAutomated safety check: PassUPL-1.01 mo ago
14

Implement a production SageMaker endpoint with autoscaling, CloudWatch alarms, and tags.

waybarrios/opencode-power-pack534—~4.6kAutomated safety check: PassApache-2.04 days ago
15

Deploy vLLM using Docker (pre-built images or build-from-source) with NVIDIA GPU support and run the OpenAI-compatible server.

vllm-project/vllm-skills102—~2.5kAutomated safety check: NotesApache-2.06 mo ago
16

Build, test, and debug Hermes Agent RL environments for Atropos training.

Tommy-yw/RunbookHermes546—~3.3kAutomated safety check: PassMIT4 mo ago
17

Configures Hyperloom after pip install --target . An agent skill from AMD-AGI/Hyperloom.

AMD-AGI/Hyperloom219—~7.1kAutomated safety check: NotesUnknowntoday
18

Deploy an API gateway for LLM traffic with load balancing, rate limiting, key management, semantic caching, fallback routing, and cost tracking.

sickn33/agentic-awesome-skills47k1 repo~2.1kAutomated safety check: PassMITyesterday
19

Auto-scale LLM inference clusters on Kubernetes using KEDA, custom GPU metrics, and horizontal pod autoscaling.

sickn33/agentic-awesome-skills47k1 repo~2.1kAutomated safety check: PassMITyesterday
20

Generates and updates secure, production-ready Kubernetes YAML manifests optimized for GKE Autopilot and GKE Standard clusters.

google/skills21k—~3.1kAutomated safety check: PassApache-2.0today
21
21.ML AIOfficial

Turn on AI + ML features in Grafana Cloud — Grafana Assistant (NL → PromQL/LogQL/TraceQL, dashboard build, incident investigation, MCP integration), Dynamic Alerting (Prophet forecasting + DBSCAN…

grafana/skills282—~1.3kAutomated safety check: PassApache-2.0yesterday
22

[omh] Self-hosted LLM serving on GPUs: choose the serving engine and quantization from decision tables, prepare deployment as an idempotent runbook with observed-only verification, and measure the…

rlaope/oh-my-hermes3.2k—~2.2kAutomated safety check: PassMITtoday
23

How to swap the VLM in the VSS Alerts Blueprint — covers RTVI-VLM microservice deployment methods, all three VLM consumers (rtvi-vlm, vlm-as-verifier, vss-agent), and health checks.

NVIDIA/skills3.6k—~5kAutomated safety check: NotesApache-2.0yesterday
24

Deploy a GPU workload on CoreWeave with kubectl. An agent skill from jeremylongshore/tons-of-skills-marketplace.

jeremylongshore/tons-of-skills-marketplace2.8k—~1.4kAutomated safety check: PassMITtoday
25

Deploy an API gateway for LLM traffic with load balancing, rate limiting, key management, semantic caching, fallback routing, and cost tracking.

BagelHole/DevOps-Security-Agent-Skills1.2k—~2kAutomated safety check: PassMIT4 mo ago
26

Auto-scale LLM inference clusters on Kubernetes using KEDA, custom GPU metrics, and horizontal pod autoscaling.

BagelHole/DevOps-Security-Agent-Skills1.2k—~2kAutomated safety check: PassMIT4 mo ago