Optimize Slurm Topology
NVlabs/alpasim
Optimize AlpaSim Slurm topology throughput using persistent local Prometheus/Grafana telemetry and run artifacts.
Set up and manage NVIDIA GPU servers for AI workloads. An agent skill from sickn33/agentic-awesome-skills.
$ npx skills add sickn33/agentic-awesome-skills --skill gpu-server-management -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install sickn33/agentic-awesome-skills gpu-server-management --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/sickn33/agentic-awesome-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/gpu-server-management .claude/skills/gpu-server-management && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "gpu-server-management" agent skill from https://github.com/sickn33/agentic-awesome-skills/tree/main/skills/gpu-server-management into .claude/skills/gpu-server-management/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gpu-server-management", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/sickn33/agentic-awesome-skills/tree/main/skills/gpu-server-managementType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add sickn33/agentic-awesome-skills --skill gpu-server-management -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install sickn33/agentic-awesome-skills gpu-server-management --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/sickn33/agentic-awesome-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/gpu-server-management .agents/skills/gpu-server-management && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "gpu-server-management" agent skill from https://github.com/sickn33/agentic-awesome-skills/tree/main/skills/gpu-server-management into .agents/skills/gpu-server-management/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gpu-server-management", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add sickn33/agentic-awesome-skills --skill gpu-server-management -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install sickn33/agentic-awesome-skills gpu-server-management --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/sickn33/agentic-awesome-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/gpu-server-management .cursor/skills/gpu-server-management && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "gpu-server-management" agent skill from https://github.com/sickn33/agentic-awesome-skills/tree/main/skills/gpu-server-management into .cursor/skills/gpu-server-management/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gpu-server-management", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/sickn33/agentic-awesome-skills.git --path skills/gpu-server-management--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add sickn33/agentic-awesome-skills --skill gpu-server-management -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install sickn33/agentic-awesome-skills gpu-server-management --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/sickn33/agentic-awesome-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/gpu-server-management .gemini/skills/gpu-server-management && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "gpu-server-management" agent skill from https://github.com/sickn33/agentic-awesome-skills/tree/main/skills/gpu-server-management into .gemini/skills/gpu-server-management/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gpu-server-management", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install sickn33/agentic-awesome-skills gpu-server-managementInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add sickn33/agentic-awesome-skills --skill gpu-server-management -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/sickn33/agentic-awesome-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/gpu-server-management .github/skills/gpu-server-management && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "gpu-server-management" agent skill from https://github.com/sickn33/agentic-awesome-skills/tree/main/skills/gpu-server-management into .github/skills/gpu-server-management/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gpu-server-management", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add sickn33/agentic-awesome-skills --skill gpu-server-management -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install sickn33/agentic-awesome-skills gpu-server-management --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/sickn33/agentic-awesome-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/gpu-server-management .opencode/skills/gpu-server-management && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "gpu-server-management" agent skill from https://github.com/sickn33/agentic-awesome-skills/tree/main/skills/gpu-server-management into .opencode/skills/gpu-server-management/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gpu-server-management", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
gpu-server-managementSet up and manage NVIDIA GPU servers for AI workloads. An agent skill from sickn33/agentic-awesome-skills.
GPU Server Management is an agent skill from sickn33/agentic-awesome-skills. Set up and manage NVIDIA GPU servers for AI workloads
Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts. Compatibility notes: Requires the relevant OS/platform tooling and privileged access where noted. Docs-only; helper scripts and templates not bundled.
It sits in AI & LLM Engineering, covering Monitoring and alerting. It works with NVIDIA AI Platform and Prometheus. The repository describes itself as: AAS Core is the local, agent-first control plane for complete catalog discovery, agent-owned selection, stack validation, and planning, backed by 2,400+ agentic skills. Includes… The licence is MIT.
Read from SKILL.md and the folder at commit ec02547. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
aptdockercurlFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
nvidia.github.ioFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Requires the relevant OS/platform tooling and privileged access where noted. Docs-only; helper scripts and templates not bundled.
From compatibility in the SKILL.md frontmatter.
GPU Server Management loads about 2k tokens when it runs. Until then it costs about 19 tokens; SKILL.md has 322 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
- Root or sudo accesssudo apt purge -y 'nvidia*' 'cuda*' 'libcuda*'sudo apt autoremove -ysudo gpg --dearmor -o /usr/share/keyrings/nvidia-container-toolkit-keyring.gpgsudo tee /etc/apt/sources.list.d/nvidia-container-toolkit.listsudo apt updatesudo apt install -y nvidia-driver-560 cuda-toolkit-12-6sudo apt install -y nvidia-container-toolkitsudo nvidia-ctk runtime configure --runtime=dockersudo systemctl restart dockerAutomated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from sickn33/agentic-awesome-skills at commit ec02547, republished under its MIT licence (© sickn33). 322 words, ~1,975 tokens.
.claude/skills/gpu-server-management/SKILL.md (or your agent's skills folder).Provision, configure, and monitor NVIDIA GPU servers for AI inference and training workloads.
Use this skill when:
# Remove old drivers
sudo apt purge -y 'nvidia*' 'cuda*' 'libcuda*'
sudo apt autoremove -y
# Add NVIDIA package repository
distribution=$(. /etc/os-release; echo $ID$VERSION_ID)
curl -fsSL https://nvidia.github.io/libnvidia-container/gpgkey | \
sudo gpg --dearmor -o /usr/share/keyrings/nvidia-container-toolkit-keyring.gpg
curl -s -L https://nvidia.github.io/libnvidia-container/$distribution/libnvidia-container.list | \
sed 's#deb https://#deb [signed-by=/usr/share/keyrings/nvidia-container-toolkit-keyring.gpg] https://#g' | \
sudo tee /etc/apt/sources.list.d/nvidia-container-toolkit.list
sudo apt update
# Install latest driver (560.x as of 2025)
sudo apt install -y nvidia-driver-560 cuda-toolkit-12-6
# Install NVIDIA Container Toolkit (Docker GPU support)
sudo apt install -y nvidia-container-toolkit
sudo nvidia-ctk runtime configure --runtime=docker
sudo systemctl restart docker
# Verify
nvidia-smi
nvcc --version
docker run --rm --gpus all nvidia/cuda:12.6.0-base-ubuntu22.04 nvidia-smi# Enable persistence mode (reduces driver initialization latency)
sudo nvidia-smi -pm 1
# Set power limit (reduce heat/noise on inference servers)
sudo nvidia-smi -pl 350 # watts; check TDP for your GPU model
# Disable ECC on inference servers (frees ~6% VRAM, less safe)
sudo nvidia-smi --ecc-config=0 # requires reboot
# Enable P2P for multi-GPU NVLink training
sudo nvidia-smi topo -m # check NVLink topology# Real-time monitoring (like htop for GPUs)
watch -n 1 nvidia-smi
# Detailed stats
nvidia-smi --query-gpu=index,name,temperature.gpu,utilization.gpu,\
utilization.memory,memory.used,memory.free,power.draw,clocks.current.graphics \
--format=csv --loop=1
# DCGM — production monitoring daemon (for clusters)
sudo apt install -y datacenter-gpu-manager
sudo systemctl start dcgm
dcgmi discovery -l # list GPUs
dcgmi diag -r 1 # quick health check
dcgmi diag -r 3 # full diagnostic (takes ~20 min)
# Check GPU errors (XID errors — important for stability)
sudo dmesg | grep -i "NVRM\|nvidia\|XID"
nvidia-smi --query-gpu=ecc.errors.corrected.volatile.total \
--format=csv,noheader# Deploy DCGM Exporter for Prometheus scraping
docker run -d \
--name dcgm-exporter \
--gpus all \
--cap-add SYS_ADMIN \
-p 9400:9400 \
--restart unless-stopped \
nvcr.io/nvidia/k8s/dcgm-exporter:latest
# Key metrics exposed:
# DCGM_FI_DEV_GPU_UTIL - GPU utilization %
# DCGM_FI_DEV_MEM_COPY_UTIL - Memory bandwidth utilization
# DCGM_FI_DEV_FB_USED - Framebuffer memory used (MB)
# DCGM_FI_DEV_SM_CLOCK - SM clock speed (MHz)
# DCGM_FI_DEV_GPU_TEMP - Temperature (°C)
# DCGM_FI_DEV_POWER_USAGE - Power draw (W)
# DCGM_FI_DEV_XID_ERRORS - XID error count (0 = healthy)MIG (Multi-Instance GPU) allows slicing one GPU into isolated smaller GPUs.
# Enable MIG mode (requires reboot or restart of all processes)
sudo nvidia-smi -mig 1
sudo systemctl restart nvidia-persistenced
# List available MIG profiles (A100 80GB example)
nvidia-smi mig -lgip
# 1g.10gb — 1 slice, 10GB (max 7 instances)
# 2g.20gb — 2 slices, 20GB (max 3 instances)
# 3g.40gb — 3 slices, 40GB (max 2 instances)
# 7g.80gb — full GPU, 80GB (max 1 instance)
# Create MIG instances (e.g., 3× 2g.20gb + 1× 2g.20gb = multi-tenant)
sudo nvidia-smi mig -cgi 2g.20gb,2g.20gb,2g.20gb,2g.20gb -C
# List created instances
nvidia-smi mig -lgi
nvidia-smi mig -lcgi
# Use in Docker
docker run --gpus '"device=MIG-GPU-xxx/0/0"' ...
# Disable MIG
sudo nvidia-smi mig -i 0 -dci
sudo nvidia-smi mig -i 0 -dgi
sudo nvidia-smi -mig 0# Increase file descriptor limits
echo '* soft nofile 1048576' | sudo tee -a /etc/security/limits.conf
echo '* hard nofile 1048576' | sudo tee -a /etc/security/limits.conf
# Disable transparent huge pages (reduces latency jitter)
echo never | sudo tee /sys/kernel/mm/transparent_hugepage/enabled
echo never | sudo tee /sys/kernel/mm/transparent_hugepage/defrag
# Persist via rc.local or systemd unit:
cat <<'EOF' | sudo tee /etc/rc.local
#!/bin/bash
echo never > /sys/kernel/mm/transparent_hugepage/enabled
echo never > /sys/kernel/mm/transparent_hugepage/defrag
nvidia-smi -pm 1
exit 0
EOF
sudo chmod +x /etc/rc.local
# PCIe performance mode
sudo nvidia-smi --auto-boost-default=0
sudo nvidia-smi --auto-boost-permission=0# Check NVLink and PCIe topology
nvidia-smi topo -m
# Output shows interconnect type:
# NV4 = NVLink 4.0 (H100 SXM)
# NV2 = NVLink 2.0 (A100 SXM)
# PHB = PCIe bus (slower; avoid for tensor parallel training)
# PIX = same PCIe switch (fast)
# Bandwidth test between GPUs
/usr/local/cuda/samples/bin/x86_64/linux/release/p2pBandwidthLatencyTest| Issue | Cause | Fix |
|---|---|---|
nvidia-smi: command not found | Driver not installed | Follow driver installation steps above |
| Driver version mismatch | CUDA/driver incompatibility | Check compatibility matrix at developer.nvidia.com |
| GPU temperature >85°C | Poor airflow or fan failure | Check nvidia-smi -q -d TEMPERATURE; reseat cooler |
| XID 79 errors | GPU hardware error | Run dcgmi diag -r 3; may need GPU replacement |
failed to open device in container | Container toolkit not configured | Run nvidia-ctk runtime configure --runtime=docker |
| Low PCIe bandwidth | Wrong slot or power limit | Check `nvidia-smi -q |
nvidia-smi -pm 1) — reduces first-request latency.vllm-server) - LLM inference on GPUsllm-fine-tuning) - GPU training setuplinux-hardening) - Secure the host OSprometheus-grafana) - Metrics dashboards© sickn33, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/gpu-server-management of sickn33/agentic-awesome-skills.
Open the folder on GitHubat commit ec02547
We found 6 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in sickn33/agentic-awesome-skills, which our catalogue first saw on October 7, 2026.
GPU Server Management next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| GPU Server Management this skillsickn33/agentic-awesome-skills | 47k | 2 repos | ~2k | Automated safety check: Notes | MIT | |
| Optimize Slurm TopologyNVlabs/alpasim | 1.3k | — | ~1.6k | Automated safety check: Pass | Apache-2.0 | |
| Doca Collectx DeploymentNVIDIA/skills | 3.5k | — | ~2.8k | Automated safety check: Pass | Apache-2.0 | |
| vLLM Model ServingOrchestra-Research/AI-Research-SKILLs | 13k | 6 repos | ~2.3k | Automated safety check: Pass | MIT | |
| DGX Spark Memory and Thermal Opswshobson/agents | 40k | 1 repos | ~2k | Automated safety check: Pass | MIT | |
| Qdrant Advisorqdrant/skills | 253 | — | ~1.7k | Automated safety check: Pass | Apache-2.0 |
NVlabs/alpasim
Optimize AlpaSim Slurm topology throughput using persistent local Prometheus/Grafana telemetry and run artifacts.
NVIDIA/skills
A skill your agent uses to deploy and operate a CollectX (clx) based DOCA telemetry collector on a host or BlueField — wiring providers / counters into the collector, running the collection daemon…
Orchestra-Research/AI-Research-SKILLs
Deploys LLMs with vLLM for high-throughput serving, covering the OpenAI-compatible server, offline batch inference, monitoring and a Docker rollout.
wshobson/agents
Plans memory headroom, works through out-of-memory failures and watches temperature and power during long ML training jobs on NVIDIA DGX Spark.
qdrant/skills
Diagnose, troubleshoot, and advise on any Qdrant deployment by loading the latest official Qdrant skills live from skills.qdrant.tech.
grafana/skills
Turn on AI + ML features in Grafana Cloud — Grafana Assistant (NL → PromQL/LogQL/TraceQL, dashboard build, incident investigation, MCP integration), Dynamic Alerting (Prophet forecasting + DBSCAN…
sickn33/agentic-awesome-skills
Implements an interface in one of two named color modes, iridescent white or colorful black, from a parameterized starter that reports measured color intensity.
sickn33/agentic-awesome-skills
Saves a user's project decisions, rules and preferences into a project-local mdbase so later sessions and other agents can recover the intent.
sickn33/agentic-awesome-skills
Keeps project decisions, research and verified results available across coding-agent sessions through LWC memory, a document Wiki graph and a CodeGraph code index.
sickn33/agentic-awesome-skills
Guides an agent through assessing its own owner for cofounder fit, publishing an approved profile, and ranking complementary profiles other agents published for their owners.
sickn33/agentic-awesome-skills
Acts as a proxy for the Cline CLI, dispatching coding tasks one at a time, monitoring runs by hard evidence, relaying decisions to you and learning per-project preferences.
sickn33/agentic-awesome-skills
Drafts and reviews audience-specific content from supplied brand examples, with local scripts for brand voice and SEO diagnostics, channel templates and a content calendar.
Works with
Categories
Set up and manage NVIDIA GPU servers for AI workloads. An agent skill from sickn33/agentic-awesome-skills. GPU Server Management is an agent skill from sickn33/agentic-awesome-skills.
GPU Server Management fits situations like: tasks that involve Monitoring and alerting.
Run `npx skills add sickn33/agentic-awesome-skills --skill gpu-server-management -a claude-code`. Or copy the skill folder (skills/gpu-server-management in sickn33/agentic-awesome-skills) into .claude/skills/gpu-server-management in your project. Claude Code loads it when a task matches its description.
Run `npx skills add sickn33/agentic-awesome-skills --skill gpu-server-management -a codex`. Or copy the skill folder (skills/gpu-server-management in sickn33/agentic-awesome-skills) into .agents/skills/gpu-server-management in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add sickn33/agentic-awesome-skills --skill gpu-server-management -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/gpu-server-management, .gemini/skills/gpu-server-management, .github/skills/gpu-server-management and .opencode/skills/gpu-server-management in your project.
Going by SKILL.md and its folder, GPU Server Management needs the command-line tools its instructions call (apt, docker and curl). Our summary lists: Docker. Compatibility (from SKILL.md): Requires the relevant OS/platform tooling and privileged access where noted. Docs-only; helper scripts and templates not bundled..
SKILL.md names 1 domain. In commands or code: nvidia.github.io; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (runs commands with sudo), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.
GPU Server Management is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 2k tokens (SKILL.md is roughly 7.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with GPU Server Management: Optimize Slurm Topology (NVlabs/alpasim, 1.3k stars), Doca Collectx Deployment (NVIDIA/skills, 3.5k stars), vLLM Model Serving (Orchestra-Research/AI-Research-SKILLs, 13k stars) and DGX Spark Memory and Thermal Ops (wshobson/agents, 40k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
sickn33 (a GitHub user) maintains it in sickn33/agentic-awesome-skills, which has 47,343 GitHub stars. The repository holds 1,354 skills in this directory. The repository was last updated on October 7, 2026.
Source: sickn33/agentic-awesome-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.