Graphsignal
graphsignal/graphsignal
Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.
GPU optimization for consumer NVIDIA GPUs (8-24GB VRAM) covering mixed precision, gradient checkpointing, XGBoost GPU, CuPy/cuDF migration, and torch.compile.
$ npx skills add Mathews-Tom/armory --skill gpu-optimizer -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install Mathews-Tom/armory gpu-optimizer --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/Mathews-Tom/armory.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/gpu-optimizer .claude/skills/gpu-optimizer && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "gpu-optimizer" agent skill from https://github.com/Mathews-Tom/armory/tree/main/skills/gpu-optimizer into .claude/skills/gpu-optimizer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gpu-optimizer", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/Mathews-Tom/armory/tree/main/skills/gpu-optimizerType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add Mathews-Tom/armory --skill gpu-optimizer -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install Mathews-Tom/armory gpu-optimizer --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Mathews-Tom/armory.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/gpu-optimizer .agents/skills/gpu-optimizer && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "gpu-optimizer" agent skill from https://github.com/Mathews-Tom/armory/tree/main/skills/gpu-optimizer into .agents/skills/gpu-optimizer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gpu-optimizer", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Mathews-Tom/armory --skill gpu-optimizer -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install Mathews-Tom/armory gpu-optimizer --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Mathews-Tom/armory.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/gpu-optimizer .cursor/skills/gpu-optimizer && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "gpu-optimizer" agent skill from https://github.com/Mathews-Tom/armory/tree/main/skills/gpu-optimizer into .cursor/skills/gpu-optimizer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gpu-optimizer", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/Mathews-Tom/armory.git --path skills/gpu-optimizer--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add Mathews-Tom/armory --skill gpu-optimizer -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install Mathews-Tom/armory gpu-optimizer --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Mathews-Tom/armory.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/gpu-optimizer .gemini/skills/gpu-optimizer && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "gpu-optimizer" agent skill from https://github.com/Mathews-Tom/armory/tree/main/skills/gpu-optimizer into .gemini/skills/gpu-optimizer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gpu-optimizer", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install Mathews-Tom/armory gpu-optimizerInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add Mathews-Tom/armory --skill gpu-optimizer -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/Mathews-Tom/armory.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/gpu-optimizer .github/skills/gpu-optimizer && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "gpu-optimizer" agent skill from https://github.com/Mathews-Tom/armory/tree/main/skills/gpu-optimizer into .github/skills/gpu-optimizer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gpu-optimizer", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Mathews-Tom/armory --skill gpu-optimizer -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install Mathews-Tom/armory gpu-optimizer --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Mathews-Tom/armory.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/gpu-optimizer .opencode/skills/gpu-optimizer && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "gpu-optimizer" agent skill from https://github.com/Mathews-Tom/armory/tree/main/skills/gpu-optimizer into .opencode/skills/gpu-optimizer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gpu-optimizer", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
gpu-optimizerGPU optimization for consumer NVIDIA GPUs (8-24GB VRAM) covering mixed precision, gradient checkpointing, XGBoost GPU, CuPy/cuDF migration, and torch.compile.
GPU Optimizer is an agent skill from Mathews-Tom/armory. GPU optimization for consumer NVIDIA GPUs (8-24GB VRAM) covering mixed precision, gradient checkpointing, XGBoost GPU, CuPy/cuDF migration, and torch.compile. Triggers on: "optimize GPU training", "speed up CUDA", "reduce OOM", "migrate NumPy to CuPy", "manage GPU memory", "benchmark PyTorch".
Its SKILL.md is about 3.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `evals/cases.yaml`).
It sits in AI & LLM Engineering, covering GPU and accelerator computing, Deep learning and Machine learning. It works with PyTorch, NumPy, NVIDIA AI Platform and CUDA. The repository describes itself as: Curated, production-grade skills for AI coding agents. Battle-tested workflows for developers who use AI seriously. The licence is MIT.
9 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 4594fb7. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
uvFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
pypi.nvidia.comAlso links to:
xgboost.readthedocs.iopytorch.orgdocs.cupy.devdocs.rapids.aiFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
GPU Optimizer loads about 3.5k tokens when it runs. Until then it costs about 77 tokens; SKILL.md has 667 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
fails, verify driver installation with `sudo nvidia-smi` or reinstall drivers before proceeding.Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from Mathews-Tom/armory at commit 4594fb7, republished under its MIT licence (© Mathews-Tom). 667 words, ~3,495 tokens.
.claude/skills/gpu-optimizer/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.Expert GPU optimization for consumer GPUs with 8–24GB VRAM. Evidence-based patterns only.
Fill in your hardware before applying optimizations:
| Property | Your Value |
|---|---|
| GPU model | (e.g., RTX 4080 Mobile, RTX 3090, RTX 4090) |
| VRAM | (e.g., 12GB, 16GB, 24GB) |
| CUDA version | (nvidia-smi → top-right) |
| TDP / power limit | (laptop vs desktop affects sustained throughput) |
| Driver version | (nvidia-smi → top-left) |
Key constraint: VRAM capacity determines which strategies apply. Patterns below are annotated with minimum VRAM requirements where relevant.
DMatrix vs QuantileDMatrix:
# GPU-optimized: QuantileDMatrix is 1.8x faster
dtrain = xgb.QuantileDMatrix(X_train.astype(np.float32))
dval = xgb.QuantileDMatrix(X_val.astype(np.float32))
# Standard: DMatrix (use for inference only)
dtest = xgb.DMatrix(X_test.astype(np.float32))Critical Parameters:
params = {
'tree_method': 'hist', # GPU-accelerated histogram
'device': 'cuda:0', # Explicit GPU device
'max_bin': 256, # Higher bins = better splits (VRAM permitting)
'grow_policy': 'depthwise', # vs 'lossguide' for imbalanced data
'predictor': 'gpu_predictor', # GPU inference
}
# Training with explicit device
model = xgb.train(params, dtrain, num_boost_round=100)GPU Verification (fail-fast):
def verify_gpu():
"""Verify XGBoost GPU availability. Raises if unavailable."""
import subprocess
try:
result = subprocess.run(["nvidia-smi"], capture_output=True, text=True)
if result.returncode != 0:
raise RuntimeError("nvidia-smi failed - no GPU available")
except FileNotFoundError:
raise RuntimeError("nvidia-smi not found - no GPU available")
build_info = xgb.build_info()
if not build_info.get("USE_CUDA"):
raise RuntimeError("XGBoost not compiled with CUDA support")Memory Management:
# Single-pass training (reuse QuantileDMatrix across slots)
dtrain = xgb.QuantileDMatrix(X_train.astype(np.float32))
for slot_idx in range(num_slots):
dtrain.set_label(y_train[:, slot_idx]) # Reuse matrix
model = xgb.train(params, dtrain, num_boost_round=100)BF16 (preferred) vs FP16:
from torch.amp import autocast, GradScaler
# Auto-detect best precision
if torch.cuda.is_bf16_supported():
amp_dtype = torch.bfloat16 # Ampere+ GPUs support BF16
else:
amp_dtype = torch.float16
# Training step
scaler = GradScaler('cuda') if amp_dtype == torch.float16 else None
with autocast('cuda', dtype=amp_dtype):
output = model(input_ids, attention_mask)
loss = criterion(output, targets)
# Backward with scaling (FP16 only)
if scaler:
scaler.scale(loss).backward()
scaler.step(optimizer)
scaler.update()
else:
loss.backward()
optimizer.step()Why BF16 > FP16:
Gradient Checkpointing:
# Saves ~40% VRAM, adds ~20% compute time
model.gradient_checkpointing_enable()
# For transformers:
model.base_model.model.gradient_checkpointing_enable()VRAM Monitoring:
import torch
torch.cuda.reset_peak_memory_stats()
# ... training ...
peak_vram_gb = torch.cuda.max_memory_allocated() / 1024**3
print(f"Peak VRAM: {peak_vram_gb:.2f} GB")
# Clear cache between experiments
torch.cuda.empty_cache()Gradient Accumulation:
# Simulate larger batch size without OOM
grad_accum_steps = max(1, target_batch_size // actual_batch_size)
for i, batch in enumerate(dataloader):
loss = model(batch) / grad_accum_steps
loss.backward()
if (i + 1) % grad_accum_steps == 0:
optimizer.step()
optimizer.zero_grad()DoE for VRAM Optimization:
EXPERIMENTS = [
{"batch_size": 2, "seq_len": 128, "grad_ckpt": True, "amp": "bf16"},
{"batch_size": 4, "seq_len": 256, "grad_ckpt": True, "amp": "bf16"},
{"batch_size": 8, "seq_len": 512, "grad_ckpt": False, "amp": "bf16"},
{"batch_size": 16, "seq_len": 256, "grad_ckpt": False, "amp": "bf16"},
]Tensor Lookups (not Python loops):
# Slow: Python loop
for i, token_id in enumerate(input_ids):
type_id = token_to_type[token_id]
embeddings[i] = type_embeddings[type_id]
# Fast: Vectorized
type_ids = token_to_type[input_ids] # Broadcast lookup
embeddings = type_embeddings[type_ids] # Single GPU kernelRegistered Buffers (persistent GPU data):
class Model(nn.Module):
def __init__(self):
super().__init__()
# Build lookup tensors once
type_ids = torch.zeros(vocab_size, dtype=torch.long)
self.register_buffer('_type_ids', type_ids) # Stays on GPU
def forward(self, input_ids):
return self._type_ids[input_ids] # Vectorized lookupBatch Operations:
# Slow: Per-sample processing
outputs = [model(x.unsqueeze(0)) for x in batch]
# Fast: Batched
outputs = model(batch) # Single forward passWhen to Use CuPy:
Migration Pattern:
import cupy as cp
import numpy as np
# NumPy (CPU)
x = np.random.randn(10000, 1000)
y = np.dot(x, x.T)
# CuPy (GPU) - SAME API
x_gpu = cp.random.randn(10000, 1000)
y_gpu = cp.dot(x_gpu, x_gpu.T)
# Transfer back if needed
y_cpu = cp.asnumpy(y_gpu)Interop with PyTorch:
# CuPy → PyTorch (zero-copy)
x_cupy = cp.random.randn(1000, 1000)
x_torch = torch.as_tensor(x_cupy, device='cuda')
# PyTorch → CuPy (zero-copy)
x_torch = torch.randn(1000, 1000, device='cuda')
x_cupy = cp.asarray(x_torch)Install:
uv pip install cupy-cuda12x # For CUDA 12.xWhen to Use cuDF:
Migration Pattern:
import cudf
import pandas as pd
# Pandas (CPU)
df = pd.read_csv('large.csv')
grouped = df.groupby('category')['value'].mean()
# cuDF (GPU) - SAME API
df_gpu = cudf.read_csv('large.csv')
grouped_gpu = df_gpu.groupby('category')['value'].mean()
# Transfer back
grouped_cpu = grouped_gpu.to_pandas()XGBoost Integration:
import cudf
import xgboost as xgb
# Load data on GPU
df = cudf.read_csv('train.csv')
X = df[feature_cols]
y = df['target']
# Create DMatrix directly from cuDF (no CPU copy)
dtrain = xgb.DMatrix(X, label=y)Install:
# RAPIDS (includes cuDF, cuML, cuGraph)
uv pip install cudf-cu12 --extra-index-url=https://pypi.nvidia.comFused Optimizer:
# Check availability
use_fused = (
torch.cuda.is_available()
and "fused" in torch.optim.AdamW.__init__.__code__.co_varnames
)
optimizer = torch.optim.AdamW(
model.parameters(),
lr=1e-3,
fused=use_fused, # Single GPU kernel (2-3x faster)
)Torch Compile:
# PyTorch 2.0+ compile
if hasattr(torch, "compile"):
model = torch.compile(model, mode="reduce-overhead")cuDNN Benchmarking:
# Auto-tune kernels (slower startup, faster training)
torch.backends.cudnn.benchmark = True
# Disable for determinism
torch.backends.cudnn.deterministic = TrueWeighted Slot Loss:
class WeightedSlotLoss(nn.Module):
def __init__(self, slot_weights):
super().__init__()
self.slot_weights = torch.tensor(slot_weights)
def forward(self, logits_list, targets):
weighted_losses = []
for i, logits in enumerate(logits_list):
loss = F.cross_entropy(logits, targets[:, i])
weighted_losses.append(loss * self.slot_weights[i])
return torch.stack(weighted_losses).sum() / self.slot_weights.sum()Focal Loss (hard example mining):
class FocalLoss(nn.Module):
def __init__(self, gamma=2.0):
super().__init__()
self.gamma = gamma
def forward(self, logits, targets):
ce_loss = F.cross_entropy(logits, targets, reduction='none')
pt = torch.exp(-ce_loss)
focal_loss = ((1 - pt) ** self.gamma) * ce_loss
return focal_loss.mean()Position Embedding Cache:
class Model(nn.Module):
def __init__(self):
super().__init__()
self._pos_cache = {} # {seq_len: positions}
def forward(self, x):
T = x.size(1)
if T not in self._pos_cache:
self._pos_cache[T] = torch.arange(T, device=x.device)
# Limit cache size
if len(self._pos_cache) > 10:
self._pos_cache.pop(next(iter(self._pos_cache)))
return self.pos_embed(self._pos_cache[T])Attention Mask Cache:
def _create_causal_mask(self, T, device):
if T not in self._mask_cache:
mask = torch.triu(torch.ones(T, T), diagonal=1).bool()
self._mask_cache[T] = mask.to(device)
return self._mask_cache[T]Check GPU Utilization:
watch -n 1 nvidia-smi # Monitor in real-timeProfile PyTorch:
with torch.profiler.profile(
activities=[torch.profiler.ProfilerActivity.GPU],
with_stack=True,
) as prof:
model(batch)
print(prof.key_averages().table(sort_by="cuda_time_total"))Bottleneck Detection:
import torch.utils.bottleneck as bottleneck
bottleneck.main(['script.py'])QuantileDMatrix, set device='cuda:0'Avoid:
.cpu() in training loop (kills GPU pipeline)torch.cuda.synchronize() unnecessarily (breaks async)Documentation:
nvidia-smi first to confirm the GPU is visible; if the command fails, verify driver installation with sudo nvidia-smi or reinstall drivers before proceeding.RuntimeError: XGBoost not compiled with CUDA support: install the CUDA build via uv pip install xgboost from a CUDA-enabled environment, or build from source with -DUSE_CUDA=ON.ImportError or version mismatch): verify CUDA toolkit version with nvcc --version and install the matching CuPy wheel (e.g., cupy-cuda12x for CUDA 12.x).--extra-index-url=https://pypi.nvidia.com) and match the cudf-cu12 suffix to your CUDA major version.torch.compile produces incorrect results or crashes: disable with model = model (no compile) to isolate; known to fail on some custom ops — fall back to eager mode for those layers.Each optimization recommendation includes a before/after code pair showing the original pattern and the GPU-optimized equivalent. Performance gain estimates are provided as ranges (e.g., "1.8x faster", "~40% VRAM reduction") based on typical consumer GPU benchmarks — actual gains depend on workload and hardware. Where a change introduces a trade-off (e.g., gradient checkpointing adds compute time), the trade-off is stated explicitly inline.
© Mathews-Tom, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 1 other file in skills/gpu-optimizer of Mathews-Tom/armory.
Open the folder on GitHubat commit 4594fb7
GPU Optimizer next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| GPU Optimizer this skillMathews-Tom/armory | 327 | — | ~3.5k | Automated safety check: Notes | MIT | |
| Graphsignalgraphsignal/graphsignal | 257 | — | ~6.2k | Automated safety check: Pass | Apache-2.0 | |
| Hyperpod Version Checkerawslabs/agent-plugins | 912 | 1 repos | ~910 | Automated safety check: Pass | Apache-2.0 | |
| Optimize For GPUK-Dense-AI/scientific-agent-skills | 48k | 1 repos | ~3.4k | Automated safety check: Pass | MIT | |
| Optimize For GPUmajiayu000/claude-skill-registry | 666 | 1 repos | ~8.5k | Automated safety check: Pass | MIT | |
| PyTorch Lightning TrainingOrchestra-Research/AI-Research-SKILLs | 13k | 7 repos | ~2.3k | Automated safety check: Pass | MIT |
graphsignal/graphsignal
Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.
awslabs/agent-plugins
Check and compare software component versions on SageMaker HyperPod cluster nodes - NVIDIA drivers, CUDA toolkit, cuDNN, NCCL, EFA, AWS OFI NCCL, GDRCopy, MPI, Neuron SDK (Trainium/Inferentia)…
K-Dense-AI/scientific-agent-skills
GPU-accelerates scientific Python on NVIDIA hardware and verifies that the result is correct and faster.
majiayu000/claude-skill-registry
GPU-accelerate Python code using CuPy, Numba CUDA, Warp, cuDF, cuML, cuGraph, KvikIO, cuCIM, cuxfilter, cuVS, cuSpatial, and RAFT.
Orchestra-Research/AI-Research-SKILLs
Shows how to organize PyTorch training with Lightning's LightningModule and Trainer, covering validation, DDP, callbacks and learning-rate scheduling.
Orchestra-Research/AI-Research-SKILLs
Scales PyTorch, TensorFlow and Hugging Face training from a single GPU to multi-node clusters with Ray Train, including Ray Tune sweeps and checkpoint recovery.
Mathews-Tom/armory
Architecture reviews across 7 dimensions (structural, scalability, enterprise readiness, performance, security, ops, data) with scored reports.
Mathews-Tom/armory
Turn concepts into static HTML visuals exported as PNG or SVG files via HTML/CSS/SVG.
Mathews-Tom/armory
A skill your agent uses when analyzing an existing video URL or local recording: "watch this video", "analyze youtube video", "summarize this video", "youtube transcript", "find this moment", "what…
Mathews-Tom/armory
Deep code simplification and refactoring preserving behavior across Python, Go, TypeScript, Rust.
Mathews-Tom/armory
Turn concepts into animated explainer videos using Manim (Python) with MP4/GIF output, audio overlay, multi-scene composition.
Mathews-Tom/armory
Maps the unresolved architecture, policy, and scope decisions that must be answered before planning can start: one durable decision ticket per question on the issue tracker, typed and blocker-linked…
Works with
Categories
GPU optimization for consumer NVIDIA GPUs (8-24GB VRAM) covering mixed precision, gradient checkpointing, XGBoost GPU, CuPy/cuDF migration, and torch.compile. GPU Optimizer is an agent skill from Mathews-Tom/armory.compile.
GPU Optimizer fits situations like: : optimize GPU training; migrate NumPy to CuPy; manage GPU memory; benchmark PyTorch.
Run `npx skills add Mathews-Tom/armory --skill gpu-optimizer -a claude-code`. Or copy the skill folder (skills/gpu-optimizer in Mathews-Tom/armory) into .claude/skills/gpu-optimizer in your project. Claude Code loads it when a task matches its description.
Run `npx skills add Mathews-Tom/armory --skill gpu-optimizer -a codex`. Or copy the skill folder (skills/gpu-optimizer in Mathews-Tom/armory) into .agents/skills/gpu-optimizer in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Mathews-Tom/armory --skill gpu-optimizer -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/gpu-optimizer, .gemini/skills/gpu-optimizer, .github/skills/gpu-optimizer and .opencode/skills/gpu-optimizer in your project.
Going by SKILL.md and its folder, GPU Optimizer needs the command-line tools its instructions call (uv). Our summary lists: Python 3.
SKILL.md names 5 domains. In commands or code: pypi.nvidia.com; the agent is likely to contact it when it follows the instructions. As links in the text: xgboost.readthedocs.io, pytorch.org, docs.cupy.dev and docs.rapids.ai. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (runs commands with sudo), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.
GPU Optimizer is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.5k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with GPU Optimizer: Graphsignal (graphsignal/graphsignal, 257 stars), Hyperpod Version Checker (awslabs/agent-plugins, 912 stars), Optimize For GPU (K-Dense-AI/scientific-agent-skills, 48k stars) and Optimize For GPU (majiayu000/claude-skill-registry, 666 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Mathews-Tom (a GitHub user) maintains it in Mathews-Tom/armory, which has 327 GitHub stars. The repository holds 80 skills in this directory. The repository was last updated on October 6, 2026.
Source: Mathews-Tom/armory on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.