Agent skill

GPU Server Management

by sickn33 in sickn33/agentic-awesome-skills

Set up and manage NVIDIA GPU servers for AI workloads. An agent skill from sickn33/agentic-awesome-skills.

MITAuto-check: notesAI & LLM Engineering

Install GPU Server Management

skills CLI
$ npx skills add sickn33/agentic-awesome-skills --skill gpu-server-management -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install sickn33/agentic-awesome-skills gpu-server-management --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/sickn33/agentic-awesome-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/gpu-server-management .claude/skills/gpu-server-management && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
gpu-server-management
GitHub stars
47k
Used in
2 other repos
Token cost
~2k tokens
SKILL.md length
322 words
Files
1
Skills in repo
1,354
Repo updated
First seen
Licence
MIT

At a glance

Set up and manage NVIDIA GPU servers for AI workloads. An agent skill from sickn33/agentic-awesome-skills.

  • Tasks that involve Monitoring and alerting
  • SKILL.md covers When to Use This Skill, Prerequisites, Driver Installation (Ubuntu) and Post-Install Configuration, plus 9 more sections
  • Calls apt, docker and curl; reaches nvidia.github.io

What it does

GPU Server Management is an agent skill from sickn33/agentic-awesome-skills. Set up and manage NVIDIA GPU servers for AI workloads

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts. Compatibility notes: Requires the relevant OS/platform tooling and privileged access where noted. Docs-only; helper scripts and templates not bundled.

It sits in AI & LLM Engineering, covering Monitoring and alerting. It works with NVIDIA AI Platform and Prometheus. The repository describes itself as: AAS Core is the local, agent-first control plane for complete catalog discovery, agent-owned selection, stack validation, and planning, backed by 2,400+ agentic skills. Includes… The licence is MIT.

When your agent uses it

  • Tasks that involve Monitoring and alerting

Example prompts

  • “/gpu-server-management”

Requirements

  • Docker
  • Compatibility (from SKILL.md): Requires the relevant OS/platform tooling and privileged access where noted. Docs-only; helper scripts and templates not bundled.

What it can do on your machine

Read from SKILL.md and the folder at commit ec02547. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • apt
    • docker
    • curl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • nvidia.github.io

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Requires the relevant OS/platform tooling and privileged access where noted. Docs-only; helper scripts and templates not bundled.

    From compatibility in the SKILL.md frontmatter.

Context cost

GPU Server Management loads about 2k tokens when it runs. Until then it costs about 19 tokens; SKILL.md has 322 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~19
When it runs · the whole SKILL.md, loaded when a task matches
~2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteRuns commands with sudoSKILL.md:36
    - Root or sudo access
  • NoteRuns commands with sudoSKILL.md:43
    sudo apt purge -y 'nvidia*' 'cuda*' 'libcuda*'
  • NoteRuns commands with sudoSKILL.md:44
    sudo apt autoremove -y
  • NoteRuns commands with sudoSKILL.md:49
    sudo gpg --dearmor -o /usr/share/keyrings/nvidia-container-toolkit-keyring.gpg
  • NoteRuns commands with sudoSKILL.md:53
    sudo tee /etc/apt/sources.list.d/nvidia-container-toolkit.list
  • NoteRuns commands with sudoSKILL.md:55
    sudo apt update
  • NoteRuns commands with sudoSKILL.md:58
    sudo apt install -y nvidia-driver-560 cuda-toolkit-12-6
  • NoteRuns commands with sudoSKILL.md:61
    sudo apt install -y nvidia-container-toolkit
  • NoteRuns commands with sudoSKILL.md:62
    sudo nvidia-ctk runtime configure --runtime=docker
  • NoteRuns commands with sudoSKILL.md:63
    sudo systemctl restart docker

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from sickn33/agentic-awesome-skills at commit ec02547, republished under its MIT licence (© sickn33). 322 words, ~1,975 tokens.

Download SKILL.mdSave it as .claude/skills/gpu-server-management/SKILL.md (or your agent's skills folder).
name
gpu-server-management
description
Set up and manage NVIDIA GPU servers for AI workloads
compatibility
Requires the relevant OS/platform tooling and privileged access where noted. Docs-only; helper scripts and templates not bundled.
category
devops
risk
critical
source
https://github.com/BagelHole/DevOps-Security-Agent-Skills
source_repo
BagelHole/DevOps-Security-Agent-Skills
source_type
community
date_added
2026-09-20
license
MIT
license_source
https://github.com/BagelHole/DevOps-Security-Agent-Skills/blob/main/LICENSE
metadata.author
devops-skills
metadata.version
1.0

GPU Server Management

Provision, configure, and monitor NVIDIA GPU servers for AI inference and training workloads.

When to Use This Skill

Use this skill when:

  • Setting up a new GPU server for LLM inference or model training
  • Installing or upgrading NVIDIA drivers and CUDA toolkit
  • Configuring Docker with NVIDIA Container Toolkit for GPU workloads
  • Partitioning A100/H100 GPUs with MIG for multi-tenant workloads
  • Troubleshooting GPU errors, driver issues, or thermal throttling

Prerequisites

  • Ubuntu 22.04 LTS (recommended) or RHEL 8/9
  • NVIDIA GPU (A10G, A100, H100, RTX 4090, or L40S recommended)
  • Root or sudo access
  • Internet access for package downloads

Driver Installation (Ubuntu)

bash
# Remove old drivers
sudo apt purge -y 'nvidia*' 'cuda*' 'libcuda*'
sudo apt autoremove -y

# Add NVIDIA package repository
distribution=$(. /etc/os-release; echo $ID$VERSION_ID)
curl -fsSL https://nvidia.github.io/libnvidia-container/gpgkey | \
  sudo gpg --dearmor -o /usr/share/keyrings/nvidia-container-toolkit-keyring.gpg

curl -s -L https://nvidia.github.io/libnvidia-container/$distribution/libnvidia-container.list | \
  sed 's#deb https://#deb [signed-by=/usr/share/keyrings/nvidia-container-toolkit-keyring.gpg] https://#g' | \
  sudo tee /etc/apt/sources.list.d/nvidia-container-toolkit.list

sudo apt update

# Install latest driver (560.x as of 2025)
sudo apt install -y nvidia-driver-560 cuda-toolkit-12-6

# Install NVIDIA Container Toolkit (Docker GPU support)
sudo apt install -y nvidia-container-toolkit
sudo nvidia-ctk runtime configure --runtime=docker
sudo systemctl restart docker

# Verify
nvidia-smi
nvcc --version
docker run --rm --gpus all nvidia/cuda:12.6.0-base-ubuntu22.04 nvidia-smi

Post-Install Configuration

bash
# Enable persistence mode (reduces driver initialization latency)
sudo nvidia-smi -pm 1

# Set power limit (reduce heat/noise on inference servers)
sudo nvidia-smi -pl 350          # watts; check TDP for your GPU model

# Disable ECC on inference servers (frees ~6% VRAM, less safe)
sudo nvidia-smi --ecc-config=0   # requires reboot

# Enable P2P for multi-GPU NVLink training
sudo nvidia-smi topo -m          # check NVLink topology

GPU Health Monitoring

bash
# Real-time monitoring (like htop for GPUs)
watch -n 1 nvidia-smi

# Detailed stats
nvidia-smi --query-gpu=index,name,temperature.gpu,utilization.gpu,\
utilization.memory,memory.used,memory.free,power.draw,clocks.current.graphics \
--format=csv --loop=1

# DCGM — production monitoring daemon (for clusters)
sudo apt install -y datacenter-gpu-manager
sudo systemctl start dcgm
dcgmi discovery -l                # list GPUs
dcgmi diag -r 1                  # quick health check
dcgmi diag -r 3                  # full diagnostic (takes ~20 min)

# Check GPU errors (XID errors — important for stability)
sudo dmesg | grep -i "NVRM\|nvidia\|XID"
nvidia-smi --query-gpu=ecc.errors.corrected.volatile.total \
  --format=csv,noheader

Prometheus GPU Metrics (DCGM Exporter)

bash
# Deploy DCGM Exporter for Prometheus scraping
docker run -d \
  --name dcgm-exporter \
  --gpus all \
  --cap-add SYS_ADMIN \
  -p 9400:9400 \
  --restart unless-stopped \
  nvcr.io/nvidia/k8s/dcgm-exporter:latest

# Key metrics exposed:
# DCGM_FI_DEV_GPU_UTIL          - GPU utilization %
# DCGM_FI_DEV_MEM_COPY_UTIL     - Memory bandwidth utilization
# DCGM_FI_DEV_FB_USED           - Framebuffer memory used (MB)
# DCGM_FI_DEV_SM_CLOCK          - SM clock speed (MHz)
# DCGM_FI_DEV_GPU_TEMP          - Temperature (°C)
# DCGM_FI_DEV_POWER_USAGE       - Power draw (W)
# DCGM_FI_DEV_XID_ERRORS        - XID error count (0 = healthy)

MIG Partitioning (A100/H100)

MIG (Multi-Instance GPU) allows slicing one GPU into isolated smaller GPUs.

bash
# Enable MIG mode (requires reboot or restart of all processes)
sudo nvidia-smi -mig 1
sudo systemctl restart nvidia-persistenced

# List available MIG profiles (A100 80GB example)
nvidia-smi mig -lgip
# 1g.10gb   — 1 slice,  10GB (max 7 instances)
# 2g.20gb   — 2 slices, 20GB (max 3 instances)
# 3g.40gb   — 3 slices, 40GB (max 2 instances)
# 7g.80gb   — full GPU, 80GB (max 1 instance)

# Create MIG instances (e.g., 3× 2g.20gb + 1× 2g.20gb = multi-tenant)
sudo nvidia-smi mig -cgi 2g.20gb,2g.20gb,2g.20gb,2g.20gb -C

# List created instances
nvidia-smi mig -lgi
nvidia-smi mig -lcgi

# Use in Docker
docker run --gpus '"device=MIG-GPU-xxx/0/0"' ...

# Disable MIG
sudo nvidia-smi mig -i 0 -dci
sudo nvidia-smi mig -i 0 -dgi
sudo nvidia-smi -mig 0

Kernel & OS Tuning for GPU Servers

bash
# Increase file descriptor limits
echo '* soft nofile 1048576' | sudo tee -a /etc/security/limits.conf
echo '* hard nofile 1048576' | sudo tee -a /etc/security/limits.conf

# Disable transparent huge pages (reduces latency jitter)
echo never | sudo tee /sys/kernel/mm/transparent_hugepage/enabled
echo never | sudo tee /sys/kernel/mm/transparent_hugepage/defrag

# Persist via rc.local or systemd unit:
cat <<'EOF' | sudo tee /etc/rc.local
#!/bin/bash
echo never > /sys/kernel/mm/transparent_hugepage/enabled
echo never > /sys/kernel/mm/transparent_hugepage/defrag
nvidia-smi -pm 1
exit 0
EOF
sudo chmod +x /etc/rc.local

# PCIe performance mode
sudo nvidia-smi --auto-boost-default=0
sudo nvidia-smi --auto-boost-permission=0

Multi-GPU Topology Check

bash
# Check NVLink and PCIe topology
nvidia-smi topo -m
# Output shows interconnect type:
# NV4 = NVLink 4.0 (H100 SXM)
# NV2 = NVLink 2.0 (A100 SXM)
# PHB = PCIe bus (slower; avoid for tensor parallel training)
# PIX = same PCIe switch (fast)

# Bandwidth test between GPUs
/usr/local/cuda/samples/bin/x86_64/linux/release/p2pBandwidthLatencyTest

Common Issues

IssueCauseFix
nvidia-smi: command not foundDriver not installedFollow driver installation steps above
Driver version mismatchCUDA/driver incompatibilityCheck compatibility matrix at developer.nvidia.com
GPU temperature >85°CPoor airflow or fan failureCheck nvidia-smi -q -d TEMPERATURE; reseat cooler
XID 79 errorsGPU hardware errorRun dcgmi diag -r 3; may need GPU replacement
failed to open device in containerContainer toolkit not configuredRun nvidia-ctk runtime configure --runtime=docker
Low PCIe bandwidthWrong slot or power limitCheck `nvidia-smi -q

Best Practices

  • Always enable persistence mode (nvidia-smi -pm 1) — reduces first-request latency.
  • Monitor XID errors; persistent XID 79/94 indicates hardware failure.
  • For training: use NVLink-connected GPUs; for inference: PCIe is usually fine.
  • Set up DCGM alerts on temperature >80°C and power draw near TDP.
  • Use MIG for multi-tenant inference to provide GPU isolation between models.
  • vllm-server (vllm-server) - LLM inference on GPUs
  • llm-fine-tuning (llm-fine-tuning) - GPU training setup
  • linux-hardening (linux-hardening) - Secure the host OS
  • prometheus-grafana (prometheus-grafana) - Metrics dashboards

Limitations

  • Infrastructure commands can disrupt services: confirm target host/scope and have backups/snapshots before mutating state.
  • Docs-only import: upstream scripts and templates not bundled.

© sickn33, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/gpu-server-management of sickn33/agentic-awesome-skills.

Open the folder on GitHubat commit ec02547

Used in 2 other repositories

We found 6 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in sickn33/agentic-awesome-skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

GPU Server Management next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

GPU Server Management compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
GPU Server Management this skillsickn33/agentic-awesome-skills47k2 repos~2kAutomated safety check: NotesMIT
Optimize Slurm TopologyNVlabs/alpasim1.3k—~1.6kAutomated safety check: PassApache-2.0
Doca Collectx DeploymentNVIDIA/skills3.5k—~2.8kAutomated safety check: PassApache-2.0
vLLM Model ServingOrchestra-Research/AI-Research-SKILLs13k6 repos~2.3kAutomated safety check: PassMIT
DGX Spark Memory and Thermal Opswshobson/agents40k1 repos~2kAutomated safety check: PassMIT
Qdrant Advisorqdrant/skills253—~1.7kAutomated safety check: PassApache-2.0

Similar skills

  • Optimize AlpaSim Slurm topology throughput using persistent local Prometheus/Grafana telemetry and run artifacts.

    1.3k GitHub stars~1.6k tokensUpdated 20 days ago
    DevOps & CloudAuto-check passed
  • Official

    A skill your agent uses to deploy and operate a CollectX (clx) based DOCA telemetry collector on a host or BlueField — wiring providers / counters into the collector, running the collection daemon…

    3.5k GitHub stars~2.8k tokensUpdated today
    DevOps & CloudAuto-check passed
  • vLLM Model Serving

    Orchestra-Research/AI-Research-SKILLs

    Deploys LLMs with vLLM for high-throughput serving, covering the OpenAI-compatible server, offline batch inference, monitoring and a Docker rollout.

    13k GitHub starsUsed in 6 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Plans memory headroom, works through out-of-memory failures and watches temperature and power during long ML training jobs on NVIDIA DGX Spark.

    40k GitHub starsUsed in 1 repo~2k tokens
    AI & LLM EngineeringAuto-check passed
  • Qdrant Advisor

    qdrant/skills

    Official

    Diagnose, troubleshoot, and advise on any Qdrant deployment by loading the latest official Qdrant skills live from skills.qdrant.tech.

    253 GitHub stars~1.7k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • ML AI

    grafana/skills

    Official

    Turn on AI + ML features in Grafana Cloud — Grafana Assistant (NL → PromQL/LogQL/TraceQL, dashboard build, incident investigation, MCP integration), Dynamic Alerting (Prophet forecasting + DBSCAN…

    279 GitHub stars~1.3k tokensUpdated yesterday
    DevOps & CloudAuto-check passed

More from sickn33/agentic-awesome-skills

All 1,354 skills in this repo
  • Liuguang Banlan UI

    sickn33/agentic-awesome-skills

    Implements an interface in one of two named color modes, iridescent white or colorful black, from a parameterized starter that reports measured color intensity.

    47k GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • User Thoughts Memory

    sickn33/agentic-awesome-skills

    Saves a user's project decisions, rules and preferences into a project-local mdbase so later sessions and other agents can recover the intent.

    47k GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • Using LWC Memory and Graphs

    sickn33/agentic-awesome-skills

    Keeps project decisions, research and verified results available across coding-agent sessions through LWC memory, a document Wiki graph and a CodeGraph code index.

    47k GitHub starsUsed in 1 repo~2k tokens
    Auto-check passed
  • Find Complementary Founders

    sickn33/agentic-awesome-skills

    Guides an agent through assessing its own owner for cofounder fit, publishing an approved profile, and ranking complementary profiles other agents published for their owners.

    47k GitHub starsUsed in 1 repo~4.8k tokens
    Auto-check passed
  • Cline Pilot

    sickn33/agentic-awesome-skills

    Acts as a proxy for the Cline CLI, dispatching coding tasks one at a time, monitoring runs by hard evidence, relaying decisions to you and learning per-project preferences.

    47k GitHub starsUsed in 1 repo~4.6k tokens
    Auto-check passed
  • Content Creator

    sickn33/agentic-awesome-skills

    Drafts and reviews audience-specific content from supplied brand examples, with local scripts for brand voice and SEO diagnostics, channel templates and a content calendar.

    47k GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed

Questions about GPU Server Management

What does GPU Server Management do?

Set up and manage NVIDIA GPU servers for AI workloads. An agent skill from sickn33/agentic-awesome-skills. GPU Server Management is an agent skill from sickn33/agentic-awesome-skills.

When should I use GPU Server Management?

GPU Server Management fits situations like: tasks that involve Monitoring and alerting.

How do I install GPU Server Management in Claude Code?

Run `npx skills add sickn33/agentic-awesome-skills --skill gpu-server-management -a claude-code`. Or copy the skill folder (skills/gpu-server-management in sickn33/agentic-awesome-skills) into .claude/skills/gpu-server-management in your project. Claude Code loads it when a task matches its description.

How do I install GPU Server Management in Codex?

Run `npx skills add sickn33/agentic-awesome-skills --skill gpu-server-management -a codex`. Or copy the skill folder (skills/gpu-server-management in sickn33/agentic-awesome-skills) into .agents/skills/gpu-server-management in your project. Codex loads it when a task matches its description.

Can I use GPU Server Management in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add sickn33/agentic-awesome-skills --skill gpu-server-management -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/gpu-server-management, .gemini/skills/gpu-server-management, .github/skills/gpu-server-management and .opencode/skills/gpu-server-management in your project.

What does GPU Server Management need to run?

Going by SKILL.md and its folder, GPU Server Management needs the command-line tools its instructions call (apt, docker and curl). Our summary lists: Docker. Compatibility (from SKILL.md): Requires the relevant OS/platform tooling and privileged access where noted. Docs-only; helper scripts and templates not bundled..

Does GPU Server Management access the network?

SKILL.md names 1 domain. In commands or code: nvidia.github.io; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is GPU Server Management safe to install?

Our automated static check of SKILL.md found notes only (runs commands with sudo), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does GPU Server Management use?

GPU Server Management is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does GPU Server Management use?

About 2k tokens (SKILL.md is roughly 7.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to GPU Server Management?

Skills that share tags, products or a category with GPU Server Management: Optimize Slurm Topology (NVlabs/alpasim, 1.3k stars), Doca Collectx Deployment (NVIDIA/skills, 3.5k stars), vLLM Model Serving (Orchestra-Research/AI-Research-SKILLs, 13k stars) and DGX Spark Memory and Thermal Ops (wshobson/agents, 40k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains GPU Server Management?

sickn33 (a GitHub user) maintains it in sickn33/agentic-awesome-skills, which has 47,343 GitHub stars. The repository holds 1,354 skills in this directory. The repository was last updated on October 7, 2026.

Source: sickn33/agentic-awesome-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.