Agent skill

Pytorch Guide

by wentorai in wentorai/research-plugins

Avoid common PyTorch mistakes and apply robust training patterns

MITAuto-check passedAI & LLM Engineering

Install Pytorch Guide

skills CLI
$ npx skills add wentorai/research-plugins --skill pytorch-guide -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install wentorai/research-plugins pytorch-guide --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/wentorai/research-plugins.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/domains/ai-ml/pytorch-guide .claude/skills/pytorch-guide && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
pytorch-guide
GitHub stars
298
Used in
1 other repo
Token cost
~2.4k tokens
SKILL.md length
321 words
Files
1
Skills in repo
405
Repo updated
First seen
Licence
MIT

At a glance

Avoid common PyTorch mistakes and apply robust training patterns

  • Tasks that involve Deep learning
  • SKILL.md covers Overview, Common Mistakes and Fixes, Robust Training Template and Performance Optimization, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Pytorch Guide is an agent skill from wentorai/research-plugins. Avoid common PyTorch mistakes and apply robust training patterns

Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering Deep learning. It works with PyTorch. The repository describes itself as: 350+ academic research skills, MCP configs, and plugins for Research-Claw and AI agents. The licence is MIT.

When your agent uses it

  • Tasks that involve Deep learning

Example prompts

  • “/pytorch-guide”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit bf44b3c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • pytorch.org
    • github.com
    • lightning.ai

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Pytorch Guide loads about 2.4k tokens when it runs. Until then it costs about 20 tokens; SKILL.md has 321 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~20
When it runs · the whole SKILL.md, loaded when a task matches
~2.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from wentorai/research-plugins at commit bf44b3c, republished under its MIT licence (© wentorai). 321 words, ~2,386 tokens.

Download SKILL.mdSave it as .claude/skills/pytorch-guide/SKILL.md (or your agent's skills folder).
name
pytorch-guide
description
Avoid common PyTorch mistakes and apply robust training patterns

PyTorch Guide

Overview

PyTorch is the dominant deep learning framework in academic research, used in the majority of papers at NeurIPS, ICML, and ICLR. Its eager execution model, Pythonic API, and seamless integration with the Python scientific stack make it the default choice for prototyping and publishing research code.

However, PyTorch's flexibility is a double-edged sword. Subtle bugs -- forgetting model.eval(), accumulating gradients across batches, incorrect device placement, memory leaks from detached tensors -- can silently corrupt results without raising errors. These issues are especially dangerous in research settings where ground truth is unknown.

This guide catalogs the most common PyTorch mistakes, provides battle-tested training patterns, and covers performance optimization techniques that every researcher should know. The patterns here are drawn from top-tier ML research codebases and the PyTorch team's own best practice recommendations.

Common Mistakes and Fixes

The Big Five Mistakes
python
# MISTAKE 1: Forgetting model.eval() and torch.no_grad()
# This causes dropout and batch norm to behave incorrectly during evaluation
# and wastes memory by tracking gradients

# WRONG
def evaluate(model, dataloader):
    total_correct = 0
    for x, y in dataloader:
        output = model(x)  # Dropout still active! BN using batch stats!
        total_correct += (output.argmax(1) == y).sum().item()

# RIGHT
@torch.no_grad()
def evaluate(model, dataloader):
    model.eval()
    total_correct = 0
    for x, y in dataloader:
        output = model(x)
        total_correct += (output.argmax(1) == y).sum().item()
    model.train()  # Restore training mode
    return total_correct
python
# MISTAKE 2: Not zeroing gradients (they accumulate by default!)
# WRONG - gradients from previous batch add to current batch
for x, y in dataloader:
    loss = criterion(model(x), y)
    loss.backward()
    optimizer.step()

# RIGHT
for x, y in dataloader:
    optimizer.zero_grad()        # Clear previous gradients
    loss = criterion(model(x), y)
    loss.backward()
    optimizer.step()

# BETTER (slightly faster, avoids memset)
for x, y in dataloader:
    optimizer.zero_grad(set_to_none=True)
    loss = criterion(model(x), y)
    loss.backward()
    optimizer.step()
python
# MISTAKE 3: Memory leaks from tensor operations in metrics
# WRONG - keeps entire computation graph in memory
losses = []
for x, y in dataloader:
    loss = criterion(model(x), y)
    losses.append(loss)  # Retains computation graph!

# RIGHT - detach from graph and move to CPU
losses = []
for x, y in dataloader:
    loss = criterion(model(x), y)
    losses.append(loss.item())  # .item() extracts Python scalar
python
# MISTAKE 4: Incorrect device placement
# WRONG - model on GPU, data on CPU
model = model.cuda()
for x, y in dataloader:
    output = model(x)  # RuntimeError: tensors on different devices

# RIGHT
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
model = model.to(device)
for x, y in dataloader:
    x, y = x.to(device), y.to(device)
    output = model(x)
python
# MISTAKE 5: Mutable default arguments in dataset transforms
# WRONG
class MyDataset(Dataset):
    def __init__(self, data, transforms=[]):  # Shared mutable list!
        self.transforms = transforms

# RIGHT
class MyDataset(Dataset):
    def __init__(self, data, transforms=None):
        self.transforms = transforms or []

Robust Training Template

python
import torch
import torch.nn as nn
from torch.utils.data import DataLoader
from torch.cuda.amp import autocast, GradScaler
import time

def train(
    model: nn.Module,
    train_loader: DataLoader,
    val_loader: DataLoader,
    optimizer: torch.optim.Optimizer,
    scheduler,
    num_epochs: int,
    device: torch.device,
    use_amp: bool = True,
):
    """Production-quality training loop with mixed precision and checkpointing."""
    criterion = nn.CrossEntropyLoss()
    scaler = GradScaler(enabled=use_amp)
    best_val_loss = float("inf")

    for epoch in range(num_epochs):
        # --- Training ---
        model.train()
        train_loss = 0.0
        t0 = time.time()

        for batch_idx, (x, y) in enumerate(train_loader):
            x, y = x.to(device, non_blocking=True), y.to(device, non_blocking=True)

            optimizer.zero_grad(set_to_none=True)

            with autocast(enabled=use_amp):
                output = model(x)
                loss = criterion(output, y)

            scaler.scale(loss).backward()
            scaler.unscale_(optimizer)
            torch.nn.utils.clip_grad_norm_(model.parameters(), max_norm=1.0)
            scaler.step(optimizer)
            scaler.update()

            train_loss += loss.item()

        scheduler.step()
        avg_train_loss = train_loss / len(train_loader)

        # --- Validation ---
        model.eval()
        val_loss = 0.0
        correct = 0
        total = 0

        with torch.no_grad():
            for x, y in val_loader:
                x, y = x.to(device, non_blocking=True), y.to(device, non_blocking=True)
                with autocast(enabled=use_amp):
                    output = model(x)
                    loss = criterion(output, y)
                val_loss += loss.item()
                correct += (output.argmax(1) == y).sum().item()
                total += y.size(0)

        avg_val_loss = val_loss / len(val_loader)
        val_acc = correct / total

        # --- Checkpoint ---
        if avg_val_loss < best_val_loss:
            best_val_loss = avg_val_loss
            torch.save({
                "epoch": epoch,
                "model_state_dict": model.state_dict(),
                "optimizer_state_dict": optimizer.state_dict(),
                "val_loss": avg_val_loss,
            }, "best_checkpoint.pt")

        elapsed = time.time() - t0
        print(f"Epoch {epoch+1}/{num_epochs} | "
              f"Train Loss: {avg_train_loss:.4f} | "
              f"Val Loss: {avg_val_loss:.4f} | "
              f"Val Acc: {val_acc:.4f} | "
              f"Time: {elapsed:.1f}s")

Performance Optimization

TechniqueSpeedupEffortWhen to Use
Mixed precision (AMP)1.5-3xLowAlways on modern GPUs
torch.compile()1.2-2xLowPyTorch 2.0+, stable models
pin_memory=True in DataLoader1.1-1.3xTrivialAlways with GPU training
non_blocking=True in .to()1.05-1.1xTrivialAlways with pinned memory
Gradient accumulationN/ALowWhen batch size limited by memory
torch.backends.cudnn.benchmark = True1.1-1.5xTrivialFixed input sizes
Distributed Data ParallelNear-linearMediumMulti-GPU training
GPU Memory Management
python
# Check GPU memory usage
print(f"Allocated: {torch.cuda.memory_allocated() / 1e9:.2f} GB")
print(f"Cached: {torch.cuda.memory_reserved() / 1e9:.2f} GB")

# Force garbage collection when debugging OOM
torch.cuda.empty_cache()
import gc; gc.collect()

# Gradient accumulation for effective large batch sizes
accumulation_steps = 4
for i, (x, y) in enumerate(dataloader):
    loss = criterion(model(x.to(device)), y.to(device)) / accumulation_steps
    loss.backward()
    if (i + 1) % accumulation_steps == 0:
        optimizer.step()
        optimizer.zero_grad(set_to_none=True)

Reproducibility Checklist

python
import torch
import numpy as np
import random

def seed_everything(seed=42):
    random.seed(seed)
    np.random.seed(seed)
    torch.manual_seed(seed)
    torch.cuda.manual_seed_all(seed)
    torch.backends.cudnn.deterministic = True
    torch.backends.cudnn.benchmark = False
    # For DataLoader workers
    def seed_worker(worker_id):
        worker_seed = seed + worker_id
        np.random.seed(worker_seed)
        random.seed(worker_seed)
    return seed_worker

seed_worker = seed_everything(42)
dataloader = DataLoader(
    dataset, batch_size=32, shuffle=True,
    worker_init_fn=seed_worker,
    generator=torch.Generator().manual_seed(42),
)

Best Practices

  • Always use torch.no_grad() for inference. It reduces memory usage by ~50%.
  • Prefer model.to(device) over .cuda(). It is device-agnostic and works on CPU, CUDA, and MPS.
  • Use torch.compile(model) on PyTorch 2.0+ for free speedups on stable architectures.
  • Profile before optimizing. Use torch.profiler to find actual bottlenecks.
  • Pin your PyTorch version in requirements.txt. Different versions can produce different numerical results.
  • Use torchinfo for model summary instead of printing the model object.

References

© wentorai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/domains/ai-ml/pytorch-guide of wentorai/research-plugins.

Open the folder on GitHubat commit bf44b3c

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in wentorai/research-plugins, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Pytorch Guide next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Pytorch Guide compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Pytorch Guide this skillwentorai/research-plugins2981 repos~2.4kAutomated safety check: PassMIT
Add Uint Supportpytorch/pytorch104k2 repos~2.3kAutomated safety check: PassCustom licence
CLIP Image-Text MatchingOrchestra-Research/AI-Research-SKILLs13k8 repos~1.7kAutomated safety check: PassMIT
Add Torch Shapes Examplefacebook/pyrefly7.1k—~1.3kAutomated safety check: PassMIT
Interview Cheatsheetwanshuiyin/ARIS-in-AI-Offer5801 repos~3.4kAutomated safety check: NotesMIT
Ghstack CIpytorch/pytorch104k—~1.4kAutomated safety check: PassCustom licence

Similar skills

  • Add Uint Support

    pytorch/pytorch

    Add unsigned integer (uint) type support to PyTorch operators by updating ATDISPATCH macros.

    104k GitHub starsUsed in 2 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed
  • CLIP Image-Text Matching

    Orchestra-Research/AI-Research-SKILLs

    Explains OpenAI's CLIP model for zero-shot image classification, image-text similarity, semantic image search and content moderation, with install steps and code patterns.

    13k GitHub starsUsed in 8 repos~1.7k tokens
    AI & LLM EngineeringAuto-check passed
  • Add Torch Shapes Example

    facebook/pyrefly

    Official

    A skill your agent uses when adding a new PyTorch model to Pyrefly's shape-tracking example corpus under tensor-shapes/pyrefly-torch-stubs/examples — i.e.

    7.1k GitHub stars~1.3k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Interview Cheatsheet

    wanshuiyin/ARIS-in-AI-Offer

    Generate a long-form Chinese interview-prep cheat sheet on a specific ML/LLM topic — formulas with derivations, from-scratch PyTorch code, comparison tables, and 25 高频面试题 (L1 必会 / L2 进阶 / L3 顶级 lab).

    580 GitHub starsUsed in 1 repo~3.4k tokens
    AI & LLM EngineeringAuto-check: notes
  • Ghstack CI

    pytorch/pytorch

    Manage CI for PyTorch ghstack stacks by running CI where its results are useful now and deferring other PRs with [no-ci].

    104k GitHub stars~1.4k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • MUSA GPU Training Optimizer

    open-infra-skills/infra-skills

    Profiles, benchmarks and tunes AI training workloads on Moore Threads MUSA GPUs with a measurement-first process that keeps model behavior unchanged.

    141 GitHub stars~1.7k tokensUpdated 3 mo ago
    AI & LLM EngineeringAuto-check passed

More from wentorai/research-plugins

All 405 skills in this repo
  • Abstract Writing Guide

    wentorai/research-plugins

    Craft structured research abstracts that maximize clarity and journal acceptance

    298 GitHub starsUsed in 1 repo~1.7k tokens
    Auto-check passed
  • Academic Citation Manager

    wentorai/research-plugins

    Manage academic citations across BibTeX, APA, MLA, and Chicago formats

    298 GitHub starsUsed in 1 repo~2.7k tokens
    Auto-check passed
  • Academic Paper Summarizer

    wentorai/research-plugins

    Summarize academic papers with structured extraction of key elements

    298 GitHub starsUsed in 1 repo~1.4k tokens
    Auto-check passed
  • Academic Study Methods

    wentorai/research-plugins

    Evidence-based study techniques for academic learning and retention

    298 GitHub starsUsed in 1 repo~1.8k tokens
    Auto-check passed
  • Academic Tone Guide

    wentorai/research-plugins

    Adjust writing tone and register for academic audiences and venues

    298 GitHub starsUsed in 1 repo~1.9k tokens
    Auto-check passed
  • Academic Translation Guide

    wentorai/research-plugins

    Academic translation, post-editing, and Chinglish correction guide

    298 GitHub starsUsed in 1 repo~1.6k tokens
    Auto-check passed

Works with

Questions about Pytorch Guide

What does Pytorch Guide do?

Avoid common PyTorch mistakes and apply robust training patterns. Pytorch Guide is an agent skill from wentorai/research-plugins.

When should I use Pytorch Guide?

Pytorch Guide fits situations like: tasks that involve Deep learning.

How do I install Pytorch Guide in Claude Code?

Run `npx skills add wentorai/research-plugins --skill pytorch-guide -a claude-code`. Or copy the skill folder (skills/domains/ai-ml/pytorch-guide in wentorai/research-plugins) into .claude/skills/pytorch-guide in your project. Claude Code loads it when a task matches its description.

How do I install Pytorch Guide in Codex?

Run `npx skills add wentorai/research-plugins --skill pytorch-guide -a codex`. Or copy the skill folder (skills/domains/ai-ml/pytorch-guide in wentorai/research-plugins) into .agents/skills/pytorch-guide in your project. Codex loads it when a task matches its description.

Can I use Pytorch Guide in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add wentorai/research-plugins --skill pytorch-guide -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/pytorch-guide, .gemini/skills/pytorch-guide, .github/skills/pytorch-guide and .opencode/skills/pytorch-guide in your project.

What does Pytorch Guide need to run?

SKILL.md names no scripts, command-line tools or credentials: Pytorch Guide is instructions for the agent only. Our summary lists: Python 3.

Does Pytorch Guide access the network?

SKILL.md names 3 domains. As links in the text: pytorch.org, github.com and lightning.ai. This is read from the text; nothing was executed.

Is Pytorch Guide safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Pytorch Guide use?

Pytorch Guide is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Pytorch Guide use?

About 2.4k tokens (SKILL.md is roughly 9.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Pytorch Guide?

Skills that share tags, products or a category with Pytorch Guide: Add Uint Support (pytorch/pytorch, 104k stars), CLIP Image-Text Matching (Orchestra-Research/AI-Research-SKILLs, 13k stars), Add Torch Shapes Example (facebook/pyrefly, 7.1k stars) and Interview Cheatsheet (wanshuiyin/ARIS-in-AI-Offer, 580 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Pytorch Guide?

wentorai (a GitHub user) maintains it in wentorai/research-plugins, which has 298 GitHub stars. The repository holds 405 skills in this directory. The repository was last updated on June 19, 2026.

Source: wentorai/research-plugins on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.