Dstack
dstackai/dstack
dstack is an open-source control plane for GPU provisioning and orchestration across GPU clouds, Kubernetes, and on-prem clusters.
Operate GPU-backed Kubernetes clusters for AI inference and training with scheduling, autoscaling, node health, MIG partitioning, and cost controls.
$ npx skills add sickn33/agentic-awesome-skills --skill gpu-kubernetes-operations -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install sickn33/agentic-awesome-skills gpu-kubernetes-operations --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/sickn33/agentic-awesome-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/gpu-kubernetes-operations .claude/skills/gpu-kubernetes-operations && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "gpu-kubernetes-operations" agent skill from https://github.com/sickn33/agentic-awesome-skills/tree/main/skills/gpu-kubernetes-operations into .claude/skills/gpu-kubernetes-operations/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gpu-kubernetes-operations", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/sickn33/agentic-awesome-skills/tree/main/skills/gpu-kubernetes-operationsType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add sickn33/agentic-awesome-skills --skill gpu-kubernetes-operations -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install sickn33/agentic-awesome-skills gpu-kubernetes-operations --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/sickn33/agentic-awesome-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/gpu-kubernetes-operations .agents/skills/gpu-kubernetes-operations && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "gpu-kubernetes-operations" agent skill from https://github.com/sickn33/agentic-awesome-skills/tree/main/skills/gpu-kubernetes-operations into .agents/skills/gpu-kubernetes-operations/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gpu-kubernetes-operations", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add sickn33/agentic-awesome-skills --skill gpu-kubernetes-operations -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install sickn33/agentic-awesome-skills gpu-kubernetes-operations --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/sickn33/agentic-awesome-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/gpu-kubernetes-operations .cursor/skills/gpu-kubernetes-operations && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "gpu-kubernetes-operations" agent skill from https://github.com/sickn33/agentic-awesome-skills/tree/main/skills/gpu-kubernetes-operations into .cursor/skills/gpu-kubernetes-operations/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gpu-kubernetes-operations", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/sickn33/agentic-awesome-skills.git --path skills/gpu-kubernetes-operations--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add sickn33/agentic-awesome-skills --skill gpu-kubernetes-operations -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install sickn33/agentic-awesome-skills gpu-kubernetes-operations --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/sickn33/agentic-awesome-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/gpu-kubernetes-operations .gemini/skills/gpu-kubernetes-operations && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "gpu-kubernetes-operations" agent skill from https://github.com/sickn33/agentic-awesome-skills/tree/main/skills/gpu-kubernetes-operations into .gemini/skills/gpu-kubernetes-operations/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gpu-kubernetes-operations", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install sickn33/agentic-awesome-skills gpu-kubernetes-operationsInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add sickn33/agentic-awesome-skills --skill gpu-kubernetes-operations -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/sickn33/agentic-awesome-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/gpu-kubernetes-operations .github/skills/gpu-kubernetes-operations && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "gpu-kubernetes-operations" agent skill from https://github.com/sickn33/agentic-awesome-skills/tree/main/skills/gpu-kubernetes-operations into .github/skills/gpu-kubernetes-operations/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gpu-kubernetes-operations", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add sickn33/agentic-awesome-skills --skill gpu-kubernetes-operations -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install sickn33/agentic-awesome-skills gpu-kubernetes-operations --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/sickn33/agentic-awesome-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/gpu-kubernetes-operations .opencode/skills/gpu-kubernetes-operations && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "gpu-kubernetes-operations" agent skill from https://github.com/sickn33/agentic-awesome-skills/tree/main/skills/gpu-kubernetes-operations into .opencode/skills/gpu-kubernetes-operations/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gpu-kubernetes-operations", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
gpu-kubernetes-operationsOperate GPU-backed Kubernetes clusters for AI inference and training with scheduling, autoscaling, node health, MIG partitioning, and cost controls.
GPU Kubernetes Operations is an agent skill from sickn33/agentic-awesome-skills. Operate GPU-backed Kubernetes clusters for AI inference and training with scheduling, autoscaling, node health, MIG partitioning, and cost controls.
Its SKILL.md is about 3.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts. Compatibility notes: Requires the relevant OS/platform tooling and privileged access where noted. Docs-only; helper scripts and templates not bundled.
It sits in DevOps & Cloud, covering Container orchestration, GPU and accelerator computing and Budgeting and forecasting. It works with Kubernetes and NVIDIA AI Platform. The repository describes itself as: AAS Core is the local, agent-first control plane for complete catalog discovery, agent-owned selection, stack validation, and planning, backed by 2,400+ agentic skills. Includes… The licence is MIT.
Read from SKILL.md and the folder at commit ec02547. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
kubectlhelmjqFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
helm.ngc.nvidia.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Requires the relevant OS/platform tooling and privileged access where noted. Docs-only; helper scripts and templates not bundled.
From compatibility in the SKILL.md frontmatter.
GPU Kubernetes Operations loads about 3.2k tokens when it runs. Until then it costs about 44 tokens; SKILL.md has 409 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from sickn33/agentic-awesome-skills at commit ec02547, republished under its MIT licence (© sickn33). 409 words, ~3,195 tokens.
.claude/skills/gpu-kubernetes-operations/SKILL.md (or your agent's skills folder).Run resilient and cost-efficient GPU clusters for production AI workloads.
The GPU Operator automates driver, toolkit, device plugin, and DCGM deployment.
# Add NVIDIA Helm repo
helm repo add nvidia https://helm.ngc.nvidia.com/nvidia
helm repo update
# Install GPU Operator
helm install gpu-operator nvidia/gpu-operator \
--namespace gpu-operator \
--create-namespace \
--set driver.enabled=true \
--set toolkit.enabled=true \
--set devicePlugin.enabled=true \
--set dcgmExporter.enabled=true \
--set migManager.enabled=true \
--set nodeStatusExporter.enabled=true \
--version v24.3.0
# Verify installation
kubectl get pods -n gpu-operator
kubectl get nodes -o json | jq '.items[].status.allocatable["nvidia.com/gpu"]'If not using the GPU Operator, deploy the device plugin directly.
# nvidia-device-plugin.yaml
apiVersion: apps/v1
kind: DaemonSet
metadata:
name: nvidia-device-plugin
namespace: kube-system
spec:
selector:
matchLabels:
name: nvidia-device-plugin
template:
metadata:
labels:
name: nvidia-device-plugin
spec:
tolerations:
- key: nvidia.com/gpu
operator: Exists
effect: NoSchedule
priorityClassName: system-node-critical
containers:
- name: nvidia-device-plugin
image: nvcr.io/nvidia/k8s-device-plugin:v0.15.0
securityContext:
privileged: true
env:
- name: FAIL_ON_INIT_ERROR
value: "false"
- name: DEVICE_SPLIT_COUNT
value: "1"
- name: DEVICE_LIST_STRATEGY
value: "envvar"
volumeMounts:
- name: device-plugin
mountPath: /var/lib/kubelet/device-plugins
volumes:
- name: device-plugin
hostPath:
path: /var/lib/kubelet/device-pluginsMIG allows a single A100 or H100 to be split into isolated GPU instances.
# mig-config.yaml - ConfigMap for MIG Manager
apiVersion: v1
kind: ConfigMap
metadata:
name: mig-parted-config
namespace: gpu-operator
data:
config.yaml: |
version: v1
mig-configs:
# 7 small instances for inference microservices
all-1g.10gb:
- devices: all
mig-enabled: true
mig-devices:
"1g.10gb": 7
# 3 medium instances for mid-size models
all-2g.20gb:
- devices: all
mig-enabled: true
mig-devices:
"2g.20gb": 3
# Mixed: 1 large + 2 small
mixed-inference:
- devices: all
mig-enabled: true
mig-devices:
"3g.40gb": 1
"1g.10gb": 4
# Full GPU for training (no partitioning)
all-disabled:
- devices: all
mig-enabled: false# Apply MIG profile to a node
kubectl label nodes gpu-node-01 nvidia.com/mig.config=all-1g.10gb --overwrite
# Verify MIG instances
kubectl exec -it nvidia-device-plugin-xxxxx -n kube-system -- nvidia-smi mig -lgi
# Check available MIG resources
kubectl get nodes gpu-node-01 -o json | jq '.status.allocatable | with_entries(select(.key | startswith("nvidia.com")))'# pod-with-mig.yaml
apiVersion: v1
kind: Pod
metadata:
name: inference-small
spec:
containers:
- name: model
image: registry.internal/vllm-server:latest
resources:
limits:
nvidia.com/mig-1g.10gb: 1
# For medium slice:
# nvidia.com/mig-2g.20gb: 1
# For large slice:
# nvidia.com/mig-3g.40gb: 1For GPUs that do not support MIG (A10, L4), use time-slicing to share a GPU.
# time-slicing-config.yaml
apiVersion: v1
kind: ConfigMap
metadata:
name: time-slicing-config
namespace: gpu-operator
data:
any: |-
version: v1
flags:
migStrategy: none
sharing:
timeSlicing:
renameByDefault: false
failRequestsGreaterThanOne: false
resources:
- name: nvidia.com/gpu
replicas: 4# Apply time-slicing config
kubectl patch clusterpolicy/cluster-policy \
--type merge \
-p '{"spec":{"devicePlugin":{"config":{"name":"time-slicing-config","default":"any"}}}}'
# After applying, each physical GPU appears as 4 virtual GPUs
kubectl get nodes -o json | jq '.items[].status.allocatable["nvidia.com/gpu"]'
# Output: "4" per physical GPU# dcgm-servicemonitor.yaml
apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata:
name: dcgm-exporter
namespace: gpu-operator
labels:
release: prometheus
spec:
selector:
matchLabels:
app: nvidia-dcgm-exporter
endpoints:
- port: gpu-metrics
interval: 15s
path: /metrics# gpu-alerts.yaml
groups:
- name: gpu-health
rules:
- alert: GPUHighTemperature
expr: DCGM_FI_DEV_GPU_TEMP > 85
for: 5m
labels:
severity: warning
annotations:
summary: "GPU {{ $labels.gpu }} temperature above 85C on {{ $labels.node }}"
- alert: GPUMemoryPressure
expr: (DCGM_FI_DEV_FB_USED / DCGM_FI_DEV_FB_FREE) > 0.90
for: 5m
labels:
severity: warning
annotations:
summary: "GPU memory above 90% on {{ $labels.node }} GPU {{ $labels.gpu }}"
- alert: GPUECCErrors
expr: increase(DCGM_FI_DEV_ECC_DBE_VOL_TOTAL[1h]) > 0
labels:
severity: critical
annotations:
summary: "Double-bit ECC errors detected on {{ $labels.node }} GPU {{ $labels.gpu }}"
- alert: GPUXidErrors
expr: increase(DCGM_FI_DEV_XID_ERRORS[5m]) > 0
labels:
severity: warning
annotations:
summary: "Xid error on {{ $labels.node }} GPU {{ $labels.gpu }}: {{ $labels.xid }}"
- alert: GPULowUtilization
expr: DCGM_FI_DEV_GPU_UTIL < 10 and on(pod) kube_pod_status_phase{phase="Running"} == 1
for: 30m
labels:
severity: info
annotations:
summary: "GPU underutilized on {{ $labels.node }} - consider rightsizing"
- alert: GPUDriverMismatch
expr: count(count by (driver_version)(DCGM_FI_DRIVER_VERSION)) > 1
labels:
severity: warning
annotations:
summary: "Multiple GPU driver versions detected across cluster"# gpu-nodepool.yaml
apiVersion: v1
kind: Node
metadata:
labels:
gpu-type: a100
gpu-memory: "80gb"
gpu-mig-capable: "true"
node-role: gpu-inference
spec:
taints:
- key: nvidia.com/gpu
value: "true"
effect: NoSchedule
---
# Inference deployment with GPU scheduling
apiVersion: apps/v1
kind: Deployment
metadata:
name: llm-inference
namespace: ai-serving
spec:
replicas: 3
selector:
matchLabels:
app: llm-inference
template:
metadata:
labels:
app: llm-inference
spec:
tolerations:
- key: nvidia.com/gpu
operator: Exists
effect: NoSchedule
nodeSelector:
gpu-type: a100
affinity:
podAntiAffinity:
preferredDuringSchedulingIgnoredDuringExecution:
- weight: 100
podAffinityTerm:
labelSelector:
matchLabels:
app: llm-inference
topologyKey: kubernetes.io/hostname
containers:
- name: vllm
image: registry.internal/vllm-server:0.4.1
resources:
requests:
nvidia.com/gpu: 1
cpu: "4"
memory: "32Gi"
limits:
nvidia.com/gpu: 1
cpu: "8"
memory: "64Gi"
env:
- name: CUDA_VISIBLE_DEVICES
value: "all"# gpu-hpa.yaml
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: llm-inference-hpa
namespace: ai-serving
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: llm-inference
minReplicas: 2
maxReplicas: 8
metrics:
- type: Pods
pods:
metric:
name: DCGM_FI_DEV_GPU_UTIL
target:
type: AverageValue
averageValue: "75"
- type: Pods
pods:
metric:
name: inference_queue_depth
target:
type: AverageValue
averageValue: "10"
behavior:
scaleUp:
stabilizationWindowSeconds: 60
policies:
- type: Pods
value: 2
periodSeconds: 120
scaleDown:
stabilizationWindowSeconds: 300
policies:
- type: Pods
value: 1
periodSeconds: 300
---
# Cluster Autoscaler config for GPU node pools
apiVersion: v1
kind: ConfigMap
metadata:
name: cluster-autoscaler-config
namespace: kube-system
data:
config: |
expander: priority
scale-down-delay-after-add: 10m
scale-down-unneeded-time: 10m
skip-nodes-with-local-storage: false
balance-similar-node-groups: true
expendable-pods-priority-cutoff: -10
gpu-total:
- min: 2
max: 16
gpu: nvidia.com/gpu| Symptom | Check | Fix |
|---|---|---|
| Pod stuck in Pending | kubectl describe pod for GPU resource events | Verify node has allocatable GPUs, check taints/tolerations |
| CUDA OOM during inference | Model too large for GPU memory | Reduce batch size, use quantization, or use MIG slice |
| DCGM metrics missing | ServiceMonitor labels matching | Verify DCGM exporter pod is running and scrape config |
| Driver mismatch after upgrade | nvidia-smi on each node | Cordon node, drain, upgrade driver, uncordon |
| GPU not detected | Device plugin pod logs | Restart device plugin, check NVIDIA container toolkit |
| Time-slicing not working | ConfigMap applied but no extra GPUs | Restart device plugin pods after config change |
| ECC errors increasing | nvidia-smi -q -d ECC | Schedule node drain and hardware replacement |
llm-inference-scaling) - Autoscale inference workloadsmodel-serving-kubernetes) - Production model serving patternsgpu-server-management) - Host-level GPU management fundamentalsmulti-tenant-llm-hosting) - Multi-tenant GPU sharingllm-cost-optimization) - Cost optimization strategies© sickn33, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/gpu-kubernetes-operations of sickn33/agentic-awesome-skills.
Open the folder on GitHubat commit ec02547
We found 6 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in sickn33/agentic-awesome-skills, which our catalogue first saw on October 7, 2026.
GPU Kubernetes Operations next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| GPU Kubernetes Operations this skillsickn33/agentic-awesome-skills | 47k | 2 repos | ~3.2k | Automated safety check: Pass | MIT | |
| Dstackdstackai/dstack | 2.3k | — | ~6.2k | Automated safety check: Warn | MPL-2.0 | |
| Paidf Orchestration SetupNVIDIA/skills | 3.5k | — | ~3.8k | Automated safety check: Warn | Apache-2.0 | |
| Nim Operator InstallNVIDIA/k8s-nim-operator | 159 | — | ~4.7k | Automated safety check: Pass | Apache-2.0 | |
| Nim Operator UninstallNVIDIA/k8s-nim-operator | 159 | — | ~3.6k | Automated safety check: Pass | Apache-2.0 | |
| Local GPU Kubernetes ValidationNVIDIA/dcgm-exporter | 1.9k | — | ~116 | Automated safety check: Pass | Apache-2.0 |
dstackai/dstack
dstack is an open-source control plane for GPU provisioning and orchestration across GPU clouds, Kubernetes, and on-prem clusters.
NVIDIA/skills
Audit, prepare, and deploy PAIDF Orchestration on a Kubernetes GPU cluster - single-GPU H100/L40S hosts, managed Kubernetes, kubeadm, and similar.
NVIDIA/k8s-nim-operator
Install NVIDIA NIM Operator on Kubernetes with prerequisite checks, optional NVIDIA GPU Operator dependency installation, public or local Helm chart selection, optional Dynamo support, and optional…
NVIDIA/k8s-nim-operator
Safely uninstall NVIDIA NIM Operator from Kubernetes with inventory checks, explicit approval gates for destructive actions, optional custom resource cleanup, optional CRD removal, and…
NVIDIA/dcgm-exporter
A skill your agent uses when validating DCGM Exporter in a local GPU-backed k3d/Kubernetes environment.
dstackai/dstack
Create and manage dstack presets: a toolkit that streamlines model inference optimization with agents, and a portable preset format.
sickn33/agentic-awesome-skills
Implements an interface in one of two named color modes, iridescent white or colorful black, from a parameterized starter that reports measured color intensity.
sickn33/agentic-awesome-skills
Saves a user's project decisions, rules and preferences into a project-local mdbase so later sessions and other agents can recover the intent.
sickn33/agentic-awesome-skills
Keeps project decisions, research and verified results available across coding-agent sessions through LWC memory, a document Wiki graph and a CodeGraph code index.
sickn33/agentic-awesome-skills
Guides an agent through assessing its own owner for cofounder fit, publishing an approved profile, and ranking complementary profiles other agents published for their owners.
sickn33/agentic-awesome-skills
Acts as a proxy for the Cline CLI, dispatching coding tasks one at a time, monitoring runs by hard evidence, relaying decisions to you and learning per-project preferences.
sickn33/agentic-awesome-skills
Drafts and reviews audience-specific content from supplied brand examples, with local scripts for brand voice and SEO diagnostics, channel templates and a content calendar.
Works with
Categories
Operate GPU-backed Kubernetes clusters for AI inference and training with scheduling, autoscaling, node health, MIG partitioning, and cost controls. GPU Kubernetes Operations is an agent skill from sickn33/agentic-awesome-skills. Operate GPU-backed Kubernetes clusters for AI inference and training with scheduling, autoscaling, node health, MIG partitioning, and cost controls.
GPU Kubernetes Operations fits situations like: tasks that involve Container orchestration; tasks that involve GPU and accelerator computing; tasks that involve Budgeting and forecasting.
Run `npx skills add sickn33/agentic-awesome-skills --skill gpu-kubernetes-operations -a claude-code`. Or copy the skill folder (skills/gpu-kubernetes-operations in sickn33/agentic-awesome-skills) into .claude/skills/gpu-kubernetes-operations in your project. Claude Code loads it when a task matches its description.
Run `npx skills add sickn33/agentic-awesome-skills --skill gpu-kubernetes-operations -a codex`. Or copy the skill folder (skills/gpu-kubernetes-operations in sickn33/agentic-awesome-skills) into .agents/skills/gpu-kubernetes-operations in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add sickn33/agentic-awesome-skills --skill gpu-kubernetes-operations -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/gpu-kubernetes-operations, .gemini/skills/gpu-kubernetes-operations, .github/skills/gpu-kubernetes-operations and .opencode/skills/gpu-kubernetes-operations in your project.
Going by SKILL.md and its folder, GPU Kubernetes Operations needs the command-line tools its instructions call (kubectl, helm and jq). Compatibility (from SKILL.md): Requires the relevant OS/platform tooling and privileged access where noted. Docs-only; helper scripts and templates not bundled..
SKILL.md names 1 domain. In commands or code: helm.ngc.nvidia.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
GPU Kubernetes Operations is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.2k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with GPU Kubernetes Operations: Dstack (dstackai/dstack, 2.3k stars), Paidf Orchestration Setup (NVIDIA/skills, 3.5k stars), Nim Operator Install (NVIDIA/k8s-nim-operator, 159 stars) and Nim Operator Uninstall (NVIDIA/k8s-nim-operator, 159 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
sickn33 (a GitHub user) maintains it in sickn33/agentic-awesome-skills, which has 47,343 GitHub stars. The repository holds 1,354 skills in this directory. The repository was last updated on October 7, 2026.
Source: sickn33/agentic-awesome-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.