Agent skill

GPU Optimizer

by Mathews-Tom in Mathews-Tom/armory

GPU optimization for consumer NVIDIA GPUs (8-24GB VRAM) covering mixed precision, gradient checkpointing, XGBoost GPU, CuPy/cuDF migration, and torch.compile.

MITAuto-check: notesAI & LLM Engineering

Install GPU Optimizer

skills CLI
$ npx skills add Mathews-Tom/armory --skill gpu-optimizer -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Mathews-Tom/armory gpu-optimizer --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Mathews-Tom/armory.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/gpu-optimizer .claude/skills/gpu-optimizer && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
gpu-optimizer
GitHub stars
327
Token cost
~3.5k tokens
SKILL.md length
667 words
Files
2
Skills in repo
80
Repo updated
First seen
Licence
MIT

At a glance

GPU optimization for consumer NVIDIA GPUs (8-24GB VRAM) covering mixed precision, gradient checkpointing, XGBoost GPU, CuPy/cuDF migration, and torch.compile.

  • Works in 9 steps: XGBoost GPU Acceleration → PyTorch Mixed Precision → VRAM Management → …
  • : optimize GPU training
  • SKILL.md covers Hardware Profile, Optimization Categories, Quick Diagnostics and Migration Checklist, plus 5 more sections
  • Calls uv; reaches pypi.nvidia.com

What it does

GPU Optimizer is an agent skill from Mathews-Tom/armory. GPU optimization for consumer NVIDIA GPUs (8-24GB VRAM) covering mixed precision, gradient checkpointing, XGBoost GPU, CuPy/cuDF migration, and torch.compile. Triggers on: "optimize GPU training", "speed up CUDA", "reduce OOM", "migrate NumPy to CuPy", "manage GPU memory", "benchmark PyTorch".

Its SKILL.md is about 3.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `evals/cases.yaml`).

It sits in AI & LLM Engineering, covering GPU and accelerator computing, Deep learning and Machine learning. It works with PyTorch, NumPy, NVIDIA AI Platform and CUDA. The repository describes itself as: Curated, production-grade skills for AI coding agents. Battle-tested workflows for developers who use AI seriously. The licence is MIT.

When your agent uses it

  • : optimize GPU training
  • Migrate NumPy to CuPy
  • Manage GPU memory
  • Benchmark PyTorch

Example prompts

  • “optimize GPU training”
  • “speed up CUDA”
  • “reduce OOM”
  • “/gpu-optimizer”

Requirements

  • Python 3

Workflow steps

9 steps, taken from the step headings in SKILL.md.

  1. XGBoost GPU Acceleration
  2. PyTorch Mixed Precision
  3. VRAM Management
  4. Aggressive Vectorization
  5. CuPy Migration (NumPy → GPU)
  6. cuDF Migration (Pandas → GPU)
  7. PyTorch Compilation & Optimization
  8. Advanced Loss Functions
  9. Caching & Precomputation

What it can do on your machine

Read from SKILL.md and the folder at commit 4594fb7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • pypi.nvidia.com

    Also links to:

    • xgboost.readthedocs.io
    • pytorch.org
    • docs.cupy.dev
    • docs.rapids.ai

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

GPU Optimizer loads about 3.5k tokens when it runs. Until then it costs about 77 tokens; SKILL.md has 667 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~77
When it runs · the whole SKILL.md, loaded when a task matches
~3.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteRuns commands with sudoSKILL.md:466
    fails, verify driver installation with `sudo nvidia-smi` or reinstall drivers before proceeding.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Mathews-Tom/armory at commit 4594fb7, republished under its MIT licence (© Mathews-Tom). 667 words, ~3,495 tokens.

Download SKILL.mdSave it as .claude/skills/gpu-optimizer/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
gpu-optimizer
description
GPU optimization for consumer NVIDIA GPUs (8-24GB VRAM) covering mixed precision, gradient checkpointing, XGBoost GPU, CuPy/cuDF migration, and torch.compile. Triggers on: "optimize GPU training", "speed up CUDA", "reduce OOM", "migrate NumPy to CuPy", "manage GPU memory", "benchmark PyTorch".
metadata.version
1.1.1
metadata.category
data
metadata.tags
gpu, cuda, vram, pytorch
metadata.difficulty
advanced
metadata.phase
build

GPU Optimizer

Expert GPU optimization for consumer GPUs with 8–24GB VRAM. Evidence-based patterns only.

Hardware Profile

Fill in your hardware before applying optimizations:

PropertyYour Value
GPU model(e.g., RTX 4080 Mobile, RTX 3090, RTX 4090)
VRAM(e.g., 12GB, 16GB, 24GB)
CUDA version(nvidia-smi → top-right)
TDP / power limit(laptop vs desktop affects sustained throughput)
Driver version(nvidia-smi → top-left)

Key constraint: VRAM capacity determines which strategies apply. Patterns below are annotated with minimum VRAM requirements where relevant.

Optimization Categories

1. XGBoost GPU Acceleration

DMatrix vs QuantileDMatrix:

python
# GPU-optimized: QuantileDMatrix is 1.8x faster
dtrain = xgb.QuantileDMatrix(X_train.astype(np.float32))
dval = xgb.QuantileDMatrix(X_val.astype(np.float32))

# Standard: DMatrix (use for inference only)
dtest = xgb.DMatrix(X_test.astype(np.float32))

Critical Parameters:

python
params = {
    'tree_method': 'hist',        # GPU-accelerated histogram
    'device': 'cuda:0',           # Explicit GPU device
    'max_bin': 256,               # Higher bins = better splits (VRAM permitting)
    'grow_policy': 'depthwise',   # vs 'lossguide' for imbalanced data
    'predictor': 'gpu_predictor', # GPU inference
}

# Training with explicit device
model = xgb.train(params, dtrain, num_boost_round=100)

GPU Verification (fail-fast):

python
def verify_gpu():
    """Verify XGBoost GPU availability. Raises if unavailable."""
    import subprocess
    try:
        result = subprocess.run(["nvidia-smi"], capture_output=True, text=True)
        if result.returncode != 0:
            raise RuntimeError("nvidia-smi failed - no GPU available")
    except FileNotFoundError:
        raise RuntimeError("nvidia-smi not found - no GPU available")

    build_info = xgb.build_info()
    if not build_info.get("USE_CUDA"):
        raise RuntimeError("XGBoost not compiled with CUDA support")

Memory Management:

python
# Single-pass training (reuse QuantileDMatrix across slots)
dtrain = xgb.QuantileDMatrix(X_train.astype(np.float32))
for slot_idx in range(num_slots):
    dtrain.set_label(y_train[:, slot_idx])  # Reuse matrix
    model = xgb.train(params, dtrain, num_boost_round=100)
2. PyTorch Mixed Precision

BF16 (preferred) vs FP16:

python
from torch.amp import autocast, GradScaler

# Auto-detect best precision
if torch.cuda.is_bf16_supported():
    amp_dtype = torch.bfloat16  # Ampere+ GPUs support BF16
else:
    amp_dtype = torch.float16

# Training step
scaler = GradScaler('cuda') if amp_dtype == torch.float16 else None

with autocast('cuda', dtype=amp_dtype):
    output = model(input_ids, attention_mask)
    loss = criterion(output, targets)

# Backward with scaling (FP16 only)
if scaler:
    scaler.scale(loss).backward()
    scaler.step(optimizer)
    scaler.update()
else:
    loss.backward()
    optimizer.step()

Why BF16 > FP16:

  • Same exponent range as FP32 (no overflow/underflow)
  • No GradScaler needed (simpler code)
  • Ampere and later GPUs have native BF16 Tensor cores
3. VRAM Management

Gradient Checkpointing:

python
# Saves ~40% VRAM, adds ~20% compute time
model.gradient_checkpointing_enable()

# For transformers:
model.base_model.model.gradient_checkpointing_enable()

VRAM Monitoring:

python
import torch

torch.cuda.reset_peak_memory_stats()
# ... training ...
peak_vram_gb = torch.cuda.max_memory_allocated() / 1024**3
print(f"Peak VRAM: {peak_vram_gb:.2f} GB")

# Clear cache between experiments
torch.cuda.empty_cache()

Gradient Accumulation:

python
# Simulate larger batch size without OOM
grad_accum_steps = max(1, target_batch_size // actual_batch_size)

for i, batch in enumerate(dataloader):
    loss = model(batch) / grad_accum_steps
    loss.backward()

    if (i + 1) % grad_accum_steps == 0:
        optimizer.step()
        optimizer.zero_grad()

DoE for VRAM Optimization:

python
EXPERIMENTS = [
    {"batch_size": 2,  "seq_len": 128, "grad_ckpt": True,  "amp": "bf16"},
    {"batch_size": 4,  "seq_len": 256, "grad_ckpt": True,  "amp": "bf16"},
    {"batch_size": 8,  "seq_len": 512, "grad_ckpt": False, "amp": "bf16"},
    {"batch_size": 16, "seq_len": 256, "grad_ckpt": False, "amp": "bf16"},
]
4. Aggressive Vectorization

Tensor Lookups (not Python loops):

python
# Slow: Python loop
for i, token_id in enumerate(input_ids):
    type_id = token_to_type[token_id]
    embeddings[i] = type_embeddings[type_id]

# Fast: Vectorized
type_ids = token_to_type[input_ids]  # Broadcast lookup
embeddings = type_embeddings[type_ids]  # Single GPU kernel

Registered Buffers (persistent GPU data):

python
class Model(nn.Module):
    def __init__(self):
        super().__init__()
        # Build lookup tensors once
        type_ids = torch.zeros(vocab_size, dtype=torch.long)
        self.register_buffer('_type_ids', type_ids)  # Stays on GPU

    def forward(self, input_ids):
        return self._type_ids[input_ids]  # Vectorized lookup

Batch Operations:

python
# Slow: Per-sample processing
outputs = [model(x.unsqueeze(0)) for x in batch]

# Fast: Batched
outputs = model(batch)  # Single forward pass
5. CuPy Migration (NumPy → GPU)

When to Use CuPy:

  • Large array operations (>1M elements)
  • Repeated NumPy calls in tight loops
  • Preprocessing pipelines before PyTorch/XGBoost

Migration Pattern:

python
import cupy as cp
import numpy as np

# NumPy (CPU)
x = np.random.randn(10000, 1000)
y = np.dot(x, x.T)

# CuPy (GPU) - SAME API
x_gpu = cp.random.randn(10000, 1000)
y_gpu = cp.dot(x_gpu, x_gpu.T)

# Transfer back if needed
y_cpu = cp.asnumpy(y_gpu)

Interop with PyTorch:

python
# CuPy → PyTorch (zero-copy)
x_cupy = cp.random.randn(1000, 1000)
x_torch = torch.as_tensor(x_cupy, device='cuda')

# PyTorch → CuPy (zero-copy)
x_torch = torch.randn(1000, 1000, device='cuda')
x_cupy = cp.asarray(x_torch)

Install:

bash
uv pip install cupy-cuda12x  # For CUDA 12.x
6. cuDF Migration (Pandas → GPU)

When to Use cuDF:

  • DataFrames >1GB
  • Groupby/aggregation on large data
  • ETL pipelines before model training

Migration Pattern:

python
import cudf
import pandas as pd

# Pandas (CPU)
df = pd.read_csv('large.csv')
grouped = df.groupby('category')['value'].mean()

# cuDF (GPU) - SAME API
df_gpu = cudf.read_csv('large.csv')
grouped_gpu = df_gpu.groupby('category')['value'].mean()

# Transfer back
grouped_cpu = grouped_gpu.to_pandas()

XGBoost Integration:

python
import cudf
import xgboost as xgb

# Load data on GPU
df = cudf.read_csv('train.csv')
X = df[feature_cols]
y = df['target']

# Create DMatrix directly from cuDF (no CPU copy)
dtrain = xgb.DMatrix(X, label=y)

Install:

bash
# RAPIDS (includes cuDF, cuML, cuGraph)
uv pip install cudf-cu12 --extra-index-url=https://pypi.nvidia.com
7. PyTorch Compilation & Optimization

Fused Optimizer:

python
# Check availability
use_fused = (
    torch.cuda.is_available()
    and "fused" in torch.optim.AdamW.__init__.__code__.co_varnames
)

optimizer = torch.optim.AdamW(
    model.parameters(),
    lr=1e-3,
    fused=use_fused,  # Single GPU kernel (2-3x faster)
)

Torch Compile:

python
# PyTorch 2.0+ compile
if hasattr(torch, "compile"):
    model = torch.compile(model, mode="reduce-overhead")

cuDNN Benchmarking:

python
# Auto-tune kernels (slower startup, faster training)
torch.backends.cudnn.benchmark = True

# Disable for determinism
torch.backends.cudnn.deterministic = True
8. Advanced Loss Functions

Weighted Slot Loss:

python
class WeightedSlotLoss(nn.Module):
    def __init__(self, slot_weights):
        super().__init__()
        self.slot_weights = torch.tensor(slot_weights)

    def forward(self, logits_list, targets):
        weighted_losses = []
        for i, logits in enumerate(logits_list):
            loss = F.cross_entropy(logits, targets[:, i])
            weighted_losses.append(loss * self.slot_weights[i])
        return torch.stack(weighted_losses).sum() / self.slot_weights.sum()

Focal Loss (hard example mining):

python
class FocalLoss(nn.Module):
    def __init__(self, gamma=2.0):
        super().__init__()
        self.gamma = gamma

    def forward(self, logits, targets):
        ce_loss = F.cross_entropy(logits, targets, reduction='none')
        pt = torch.exp(-ce_loss)
        focal_loss = ((1 - pt) ** self.gamma) * ce_loss
        return focal_loss.mean()
9. Caching & Precomputation

Position Embedding Cache:

python
class Model(nn.Module):
    def __init__(self):
        super().__init__()
        self._pos_cache = {}  # {seq_len: positions}

    def forward(self, x):
        T = x.size(1)
        if T not in self._pos_cache:
            self._pos_cache[T] = torch.arange(T, device=x.device)
            # Limit cache size
            if len(self._pos_cache) > 10:
                self._pos_cache.pop(next(iter(self._pos_cache)))
        return self.pos_embed(self._pos_cache[T])

Attention Mask Cache:

python
def _create_causal_mask(self, T, device):
    if T not in self._mask_cache:
        mask = torch.triu(torch.ones(T, T), diagonal=1).bool()
        self._mask_cache[T] = mask.to(device)
    return self._mask_cache[T]

Quick Diagnostics

Check GPU Utilization:

bash
watch -n 1 nvidia-smi  # Monitor in real-time

Profile PyTorch:

python
with torch.profiler.profile(
    activities=[torch.profiler.ProfilerActivity.GPU],
    with_stack=True,
) as prof:
    model(batch)

print(prof.key_averages().table(sort_by="cuda_time_total"))

Bottleneck Detection:

python
import torch.utils.bottleneck as bottleneck
bottleneck.main(['script.py'])

Migration Checklist

  • XGBoost: Use QuantileDMatrix, set device='cuda:0'
  • PyTorch: Enable BF16/FP16, fused optimizer, torch.compile
  • VRAM: Gradient checkpointing if approaching VRAM limit
  • NumPy→CuPy: For preprocessing >1M elements
  • Pandas→cuDF: For DataFrames >1GB
  • Vectorization: Replace Python loops with tensor ops
  • Caching: Precompute positions, masks, embeddings
  • Monitor: Track VRAM usage, profile GPU kernels

Anti-Patterns

Avoid:

  • Using .cpu() in training loop (kills GPU pipeline)
  • Creating tensors on CPU then moving to GPU (create on GPU directly)
  • Using Python loops over tensors (vectorize)
  • Ignoring VRAM monitoring (leads to OOM crashes)
  • Using FP32 when BF16/FP16 works (wastes bandwidth)
  • Calling torch.cuda.synchronize() unnecessarily (breaks async)

References

Documentation:

Show full SKILL.md (313 more words)Show less

Error Handling

  • CUDA not available at runtime: run nvidia-smi first to confirm the GPU is visible; if the command fails, verify driver installation with sudo nvidia-smi or reinstall drivers before proceeding.
  • XGBoost raises RuntimeError: XGBoost not compiled with CUDA support: install the CUDA build via uv pip install xgboost from a CUDA-enabled environment, or build from source with -DUSE_CUDA=ON.
  • OOM during training: reduce batch size first (halve it), then enable gradient checkpointing; if OOM persists after both, enable gradient accumulation to simulate the original batch size.
  • CuPy import failure (ImportError or version mismatch): verify CUDA toolkit version with nvcc --version and install the matching CuPy wheel (e.g., cupy-cuda12x for CUDA 12.x).
  • cuDF install fails or produces CUDA version errors: use the NVIDIA PyPI index (--extra-index-url=https://pypi.nvidia.com) and match the cudf-cu12 suffix to your CUDA major version.
  • torch.compile produces incorrect results or crashes: disable with model = model (no compile) to isolate; known to fail on some custom ops — fall back to eager mode for those layers.

Limitations

  • NVIDIA GPUs only — AMD (ROCm) and Intel Arc GPUs are not covered by these patterns.
  • Assumes a single-GPU setup; multi-GPU (DDP, FSDP) requires additional configuration not covered here.
  • Patterns are calibrated for consumer GPUs (8–24GB VRAM); datacenter GPUs (A100, H100) have different memory hierarchies and may benefit from different strategies.
  • Framework coverage: PyTorch, XGBoost, and RAPIDS (CuPy/cuDF) only — JAX, TensorFlow, and MXNet are out of scope.
  • Laptop GPU TDP limits sustained throughput; power-throttled performance can differ significantly from desktop benchmarks even at the same VRAM capacity.

Output Format

Each optimization recommendation includes a before/after code pair showing the original pattern and the GPU-optimized equivalent. Performance gain estimates are provided as ranges (e.g., "1.8x faster", "~40% VRAM reduction") based on typical consumer GPU benchmarks — actual gains depend on workload and hardware. Where a change introduces a trade-off (e.g., gradient checkpointing adds compute time), the trade-off is stated explicitly inline.

© Mathews-Tom, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in skills/gpu-optimizer of Mathews-Tom/armory.

  • SKILL.md
  • evals/cases.yaml

Open the folder on GitHubat commit 4594fb7

Compare with similar skills

GPU Optimizer next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

GPU Optimizer compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
GPU Optimizer this skillMathews-Tom/armory327—~3.5kAutomated safety check: NotesMIT
Graphsignalgraphsignal/graphsignal257—~6.2kAutomated safety check: PassApache-2.0
Hyperpod Version Checkerawslabs/agent-plugins9121 repos~910Automated safety check: PassApache-2.0
Optimize For GPUK-Dense-AI/scientific-agent-skills48k1 repos~3.4kAutomated safety check: PassMIT
Optimize For GPUmajiayu000/claude-skill-registry6661 repos~8.5kAutomated safety check: PassMIT
PyTorch Lightning TrainingOrchestra-Research/AI-Research-SKILLs13k7 repos~2.3kAutomated safety check: PassMIT

Similar skills

  • Graphsignal

    graphsignal/graphsignal

    Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

    257 GitHub stars~6.2k tokensUpdated 9 days ago
    AI & LLM EngineeringAuto-check passed
  • Hyperpod Version Checker

    awslabs/agent-plugins

    Official

    Check and compare software component versions on SageMaker HyperPod cluster nodes - NVIDIA drivers, CUDA toolkit, cuDNN, NCCL, EFA, AWS OFI NCCL, GDRCopy, MPI, Neuron SDK (Trainium/Inferentia)…

    912 GitHub starsUsed in 1 repo~910 tokens
    AI & LLM EngineeringAuto-check passed
  • Optimize For GPU

    K-Dense-AI/scientific-agent-skills

    GPU-accelerates scientific Python on NVIDIA hardware and verifies that the result is correct and faster.

    48k GitHub starsUsed in 1 repo~3.4k tokens
    Data & AnalyticsAuto-check passed
  • Optimize For GPU

    majiayu000/claude-skill-registry

    GPU-accelerate Python code using CuPy, Numba CUDA, Warp, cuDF, cuML, cuGraph, KvikIO, cuCIM, cuxfilter, cuVS, cuSpatial, and RAFT.

    666 GitHub starsUsed in 1 repo~8.5k tokens
    Data & AnalyticsAuto-check passed
  • PyTorch Lightning Training

    Orchestra-Research/AI-Research-SKILLs

    Shows how to organize PyTorch training with Lightning's LightningModule and Trainer, covering validation, DDP, callbacks and learning-rate scheduling.

    13k GitHub starsUsed in 7 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Ray Train Distributed Training

    Orchestra-Research/AI-Research-SKILLs

    Scales PyTorch, TensorFlow and Hugging Face training from a single GPU to multi-node clusters with Ray Train, including Ray Tune sweeps and checkpoint recovery.

    13k GitHub starsUsed in 3 repos~2.7k tokens
    AI & LLM EngineeringAuto-check passed

More from Mathews-Tom/armory

All 80 skills in this repo
  • Architecture Reviewer

    Mathews-Tom/armory

    Architecture reviews across 7 dimensions (structural, scalability, enterprise readiness, performance, security, ops, data) with scored reports.

    327 GitHub stars~4.6k tokensUpdated yesterday
    Auto-check passed
  • Concept To Image

    Mathews-Tom/armory

    Turn concepts into static HTML visuals exported as PNG or SVG files via HTML/CSS/SVG.

    327 GitHub stars~2.6k tokensUpdated yesterday
    Auto-check passed
  • Watch

    Mathews-Tom/armory

    A skill your agent uses when analyzing an existing video URL or local recording: "watch this video", "analyze youtube video", "summarize this video", "youtube transcript", "find this moment", "what…

    327 GitHub stars~2.8k tokensUpdated yesterday
    Auto-check passed
  • Code Refiner

    Mathews-Tom/armory

    Deep code simplification and refactoring preserving behavior across Python, Go, TypeScript, Rust.

    327 GitHub stars~3.1k tokensUpdated yesterday
    Auto-check passed
  • Concept To Video

    Mathews-Tom/armory

    Turn concepts into animated explainer videos using Manim (Python) with MP4/GIF output, audio overlay, multi-scene composition.

    327 GitHub stars~4.9k tokensUpdated yesterday
    Auto-check passed
  • Decision Map

    Mathews-Tom/armory

    Maps the unresolved architecture, policy, and scope decisions that must be answered before planning can start: one durable decision ticket per question on the issue tracker, typed and blocker-linked…

    327 GitHub stars~2.7k tokensUpdated yesterday
    Auto-check passed

Questions about GPU Optimizer

What does GPU Optimizer do?

GPU optimization for consumer NVIDIA GPUs (8-24GB VRAM) covering mixed precision, gradient checkpointing, XGBoost GPU, CuPy/cuDF migration, and torch.compile. GPU Optimizer is an agent skill from Mathews-Tom/armory.compile.

When should I use GPU Optimizer?

GPU Optimizer fits situations like: : optimize GPU training; migrate NumPy to CuPy; manage GPU memory; benchmark PyTorch.

How do I install GPU Optimizer in Claude Code?

Run `npx skills add Mathews-Tom/armory --skill gpu-optimizer -a claude-code`. Or copy the skill folder (skills/gpu-optimizer in Mathews-Tom/armory) into .claude/skills/gpu-optimizer in your project. Claude Code loads it when a task matches its description.

How do I install GPU Optimizer in Codex?

Run `npx skills add Mathews-Tom/armory --skill gpu-optimizer -a codex`. Or copy the skill folder (skills/gpu-optimizer in Mathews-Tom/armory) into .agents/skills/gpu-optimizer in your project. Codex loads it when a task matches its description.

Can I use GPU Optimizer in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Mathews-Tom/armory --skill gpu-optimizer -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/gpu-optimizer, .gemini/skills/gpu-optimizer, .github/skills/gpu-optimizer and .opencode/skills/gpu-optimizer in your project.

What does GPU Optimizer need to run?

Going by SKILL.md and its folder, GPU Optimizer needs the command-line tools its instructions call (uv). Our summary lists: Python 3.

Does GPU Optimizer access the network?

SKILL.md names 5 domains. In commands or code: pypi.nvidia.com; the agent is likely to contact it when it follows the instructions. As links in the text: xgboost.readthedocs.io, pytorch.org, docs.cupy.dev and docs.rapids.ai. This is read from the text; nothing was executed.

Is GPU Optimizer safe to install?

Our automated static check of SKILL.md found notes only (runs commands with sudo), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does GPU Optimizer use?

GPU Optimizer is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does GPU Optimizer use?

About 3.5k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to GPU Optimizer?

Skills that share tags, products or a category with GPU Optimizer: Graphsignal (graphsignal/graphsignal, 257 stars), Hyperpod Version Checker (awslabs/agent-plugins, 912 stars), Optimize For GPU (K-Dense-AI/scientific-agent-skills, 48k stars) and Optimize For GPU (majiayu000/claude-skill-registry, 666 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains GPU Optimizer?

Mathews-Tom (a GitHub user) maintains it in Mathews-Tom/armory, which has 327 GitHub stars. The repository holds 80 skills in this directory. The repository was last updated on October 6, 2026.

Source: Mathews-Tom/armory on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.