Agent skill

PyTorch Lightning Training Setup

by davila7 in davila7/claude-code-templates

Organizes PyTorch training code into LightningModules, DataModules and Trainers, with multi-GPU strategies, callbacks and logging configured.

MITAuto-check passedAI & LLM Engineering

Install PyTorch Lightning Training Setup

skills CLI
$ npx skills add davila7/claude-code-templates --skill pytorch-lightning -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install davila7/claude-code-templates pytorch-lightning --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/davila7/claude-code-templates.git skills-src && mkdir -p .claude/skills && cp -r skills-src/cli-tool/components/skills/scientific/pytorch-lightning .claude/skills/pytorch-lightning && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
pytorch-lightning
GitHub stars
32k
Used in
14 other repos
Token cost
~1.7k tokens
SKILL.md length
551 words
Files
11 (incl. scripts, references)
Skills in repo
477
Repo updated
First seen
Licence
MIT

At a glance

Organizes PyTorch training code into LightningModules, DataModules and Trainers, with multi-GPU strategies, callbacks and logging configured.

  • Works in 7 steps: LightningModule - Model Definition → Trainer - Training Automation → LightningDataModule - Data Pipeline… → …
  • Restructuring a plain PyTorch training loop into a LightningModule
  • SKILL.md covers Overview, When to Use This Skill, Core Capabilities and Quick Workflow, plus 1 more section
  • Runs Python scripts from its folder

What it does

PyTorch Lightning cuts training boilerplate while keeping control of the model. The skill shows the agent how to split a model into a LightningModule's six sections (initialization, training, validation, test and prediction steps, and optimizer configuration), and how to wrap data handling in a LightningDataModule with prepare_data, setup and the train, validation and test dataloaders.

The Trainer handles device management, mixed precision, gradient accumulation and clipping, checkpointing, early stopping and progress bars, and can run on several GPUs or TPUs with DDP, FSDP or DeepSpeed strategies. Callbacks and logging, including W&B and TensorBoard, are covered in reference notes. Three template scripts, `scripts/template_lightning_module.py`, `scripts/template_datamodule.py` and `scripts/quick_trainer_setup.py`, give starting points.

When your agent uses it

  • Restructuring a plain PyTorch training loop into a LightningModule
  • Configuring a Trainer for multi-GPU or TPU training
  • Putting dataset loading and splits into a reusable DataModule
  • Adding checkpointing, early stopping and logging to a training run

Example prompts

  • “Convert my train.py loop into a LightningModule and a Trainer with checkpointing.”
  • “Set up a Trainer that uses DDP across four GPUs with mixed precision.”
  • “Write a LightningDataModule for my image folders with train, validation and test splits.”

Requirements

  • PyTorch and PyTorch Lightning installed
  • Multiple GPUs or TPUs for distributed runs (optional)

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. LightningModule - Model Definition
  2. Trainer - Training Automation
  3. LightningDataModule - Data Pipeline Organization
  4. Callbacks - Extensible Training Logic
  5. Logging - Experiment Tracking
  6. Distributed Training - Scale to Multiple Devices
  7. Best Practices

What it can do on your machine

Read from SKILL.md and the folder at commit 14680ec. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 3 files in scripts/ (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

PyTorch Lightning Training Setup loads about 1.7k tokens when it runs, and up to ~27k if it reads all its reference files. Until then it costs about 73 tokens; SKILL.md has 551 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~73
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~27k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from davila7/claude-code-templates at commit 14680ec, republished under its MIT licence (© davila7). 551 words, ~1,652 tokens.

Download SKILL.mdSave it as .claude/skills/pytorch-lightning/SKILL.md (or your agent's skills folder). This skill also uses 10 other files; get the full folder from GitHub.
name
pytorch-lightning
description
Deep learning framework (PyTorch Lightning). Organize PyTorch code into LightningModules, configure Trainers for multi-GPU/TPU, implement data pipelines, callbacks, logging (W&B, TensorBoard), distributed training (DDP, FSDP, DeepSpeed), for scalable neural network training.

PyTorch Lightning

Overview

PyTorch Lightning is a deep learning framework that organizes PyTorch code to eliminate boilerplate while maintaining full flexibility. Automate training workflows, multi-device orchestration, and implement best practices for neural network training and scaling across multiple GPUs/TPUs.

When to Use This Skill

This skill should be used when:

  • Building, training, or deploying neural networks using PyTorch Lightning
  • Organizing PyTorch code into LightningModules
  • Configuring Trainers for multi-GPU/TPU training
  • Implementing data pipelines with LightningDataModules
  • Working with callbacks, logging, and distributed training strategies (DDP, FSDP, DeepSpeed)
  • Structuring deep learning projects professionally

Core Capabilities

1. LightningModule - Model Definition

Organize PyTorch models into six logical sections:

  1. Initialization - __init__() and setup()
  2. Training Loop - training_step(batch, batch_idx)
  3. Validation Loop - validation_step(batch, batch_idx)
  4. Test Loop - test_step(batch, batch_idx)
  5. Prediction - predict_step(batch, batch_idx)
  6. Optimizer Configuration - configure_optimizers()

Quick template reference: See scripts/template_lightning_module.py for a complete boilerplate.

Detailed documentation: Read references/lightning_module.md for comprehensive method documentation, hooks, properties, and best practices.

2. Trainer - Training Automation

The Trainer automates the training loop, device management, gradient operations, and callbacks. Key features:

  • Multi-GPU/TPU support with strategy selection (DDP, FSDP, DeepSpeed)
  • Automatic mixed precision training
  • Gradient accumulation and clipping
  • Checkpointing and early stopping
  • Progress bars and logging

Quick setup reference: See scripts/quick_trainer_setup.py for common Trainer configurations.

Detailed documentation: Read references/trainer.md for all parameters, methods, and configuration options.

3. LightningDataModule - Data Pipeline Organization

Encapsulate all data processing steps in a reusable class:

  1. prepare_data() - Download and process data (single-process)
  2. setup() - Create datasets and apply transforms (per-GPU)
  3. train_dataloader() - Return training DataLoader
  4. val_dataloader() - Return validation DataLoader
  5. test_dataloader() - Return test DataLoader

Quick template reference: See scripts/template_datamodule.py for a complete boilerplate.

Detailed documentation: Read references/data_module.md for method details and usage patterns.

4. Callbacks - Extensible Training Logic

Add custom functionality at specific training hooks without modifying your LightningModule. Built-in callbacks include:

  • ModelCheckpoint - Save best/latest models
  • EarlyStopping - Stop when metrics plateau
  • LearningRateMonitor - Track LR scheduler changes
  • BatchSizeFinder - Auto-determine optimal batch size

Detailed documentation: Read references/callbacks.md for built-in callbacks and custom callback creation.

5. Logging - Experiment Tracking

Integrate with multiple logging platforms:

  • TensorBoard (default)
  • Weights & Biases (WandbLogger)
  • MLflow (MLFlowLogger)
  • Neptune (NeptuneLogger)
  • Comet (CometLogger)
  • CSV (CSVLogger)

Log metrics using self.log("metric_name", value) in any LightningModule method.

Detailed documentation: Read references/logging.md for logger setup and configuration.

Show full SKILL.md (237 more words)Show less
6. Distributed Training - Scale to Multiple Devices

Choose the right strategy based on model size:

  • DDP - For models <500M parameters (ResNet, smaller transformers)
  • FSDP - For models 500M+ parameters (large transformers, recommended for Lightning users)
  • DeepSpeed - For cutting-edge features and fine-grained control

Configure with: Trainer(strategy="ddp", accelerator="gpu", devices=4)

Detailed documentation: Read references/distributed_training.md for strategy comparison and configuration.

7. Best Practices
  • Device agnostic code - Use self.device instead of .cuda()
  • Hyperparameter saving - Use self.save_hyperparameters() in __init__()
  • Metric logging - Use self.log() for automatic aggregation across devices
  • Reproducibility - Use seed_everything() and Trainer(deterministic=True)
  • Debugging - Use Trainer(fast_dev_run=True) to test with 1 batch

Detailed documentation: Read references/best_practices.md for common patterns and pitfalls.

Quick Workflow

  1. Define model:

    python
    class MyModel(L.LightningModule):
        def __init__(self):
            super().__init__()
            self.save_hyperparameters()
            self.model = YourNetwork()
    
        def training_step(self, batch, batch_idx):
            x, y = batch
            loss = F.cross_entropy(self.model(x), y)
            self.log("train_loss", loss)
            return loss
    
        def configure_optimizers(self):
            return torch.optim.Adam(self.parameters())
  2. Prepare data:

    python
    # Option 1: Direct DataLoaders
    train_loader = DataLoader(train_dataset, batch_size=32)
    
    # Option 2: LightningDataModule (recommended for reusability)
    dm = MyDataModule(batch_size=32)
  3. Train:

    python
    trainer = L.Trainer(max_epochs=10, accelerator="gpu", devices=2)
    trainer.fit(model, train_loader)  # or trainer.fit(model, datamodule=dm)

Resources

scripts/

Executable Python templates for common PyTorch Lightning patterns:

  • template_lightning_module.py - Complete LightningModule boilerplate
  • template_datamodule.py - Complete LightningDataModule boilerplate
  • quick_trainer_setup.py - Common Trainer configuration examples
references/

Detailed documentation for each PyTorch Lightning component:

  • lightning_module.md - Comprehensive LightningModule guide (methods, hooks, properties)
  • trainer.md - Trainer configuration and parameters
  • data_module.md - LightningDataModule patterns and methods
  • callbacks.md - Built-in and custom callbacks
  • logging.md - Logger integrations and usage
  • distributed_training.md - DDP, FSDP, DeepSpeed comparison and setup
  • best_practices.md - Common patterns, tips, and pitfalls

© davila7, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 10 other files (scripts, references) in cli-tool/components/skills/scientific/pytorch-lightning of davila7/claude-code-templates.

  • SKILL.md
  • references/best_practices.md
  • references/callbacks.md
  • references/data_module.md
  • references/distributed_training.md
  • references/lightning_module.md
  • references/logging.md
  • references/trainer.md
  • scripts/quick_trainer_setup.py
  • scripts/template_datamodule.py
  • scripts/template_lightning_module.py

Open the folder on GitHubat commit 14680ec

Used in 14 other repositories

We found 22 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 14 other GitHub owners. This page covers the copy in davila7/claude-code-templates, which our catalogue first saw on October 7, 2026.

Compare with similar skills

PyTorch Lightning Training Setup next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

PyTorch Lightning Training Setup compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
PyTorch Lightning Training Setup this skilldavila7/claude-code-templates32k14 repos~1.7kAutomated safety check: PassMIT
PyTorch Lightning TrainingOrchestra-Research/AI-Research-SKILLs13k7 repos~2.3kAutomated safety check: PassMIT
Ray Train Distributed TrainingOrchestra-Research/AI-Research-SKILLs13k3 repos~2.7kAutomated safety check: PassMIT
GPU OptimizerMathews-Tom/armory328—~3.5kAutomated safety check: NotesMIT
Cuda Index Widthpytorch/pytorch104k—~1.6kAutomated safety check: PassCustom licence
Liger Kernel Devlinkedin/Liger-Kernel6.6k—~799Automated safety check: PassBSD-2-Clause

Similar skills

  • PyTorch Lightning Training

    Orchestra-Research/AI-Research-SKILLs

    Shows how to organize PyTorch training with Lightning's LightningModule and Trainer, covering validation, DDP, callbacks and learning-rate scheduling.

    13k GitHub starsUsed in 7 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Ray Train Distributed Training

    Orchestra-Research/AI-Research-SKILLs

    Scales PyTorch, TensorFlow and Hugging Face training from a single GPU to multi-node clusters with Ray Train, including Ray Tune sweeps and checkpoint recovery.

    13k GitHub starsUsed in 3 repos~2.7k tokens
    AI & LLM EngineeringAuto-check passed
  • GPU Optimizer

    Mathews-Tom/armory

    GPU optimization for consumer NVIDIA GPUs (8-24GB VRAM) covering mixed precision, gradient checkpointing, XGBoost GPU, CuPy/cuDF migration, and torch.compile.

    328 GitHub stars~3.5k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check: notes
  • Cuda Index Width

    pytorch/pytorch

    Choose 32-bit vs 64-bit index math in PyTorch CUDA kernels. An agent skill from pytorch/pytorch.

    104k GitHub stars~1.6k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Liger Kernel Dev

    linkedin/Liger-Kernel

    Develops production-ready Triton kernels for Liger Kernel. An agent skill from linkedin/Liger-Kernel.

    6.6k GitHub stars~799 tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • MUSA GPU Training Optimizer

    open-infra-skills/infra-skills

    Profiles, benchmarks and tunes AI training workloads on Moore Threads MUSA GPUs with a measurement-first process that keeps model behavior unchanged.

    141 GitHub stars~1.7k tokensUpdated 3 mo ago
    AI & LLM EngineeringAuto-check passed

More from davila7/claude-code-templates

All 477 skills in this repo
  • Perplexity Web Search

    davila7/claude-code-templates

    Runs web-grounded searches through Perplexity's Sonar models over OpenRouter for current events, recent literature and cited facts beyond the model's training cutoff.

    32k GitHub starsUsed in 12 repos~3.5k tokens
    Auto-check: notes
  • Neuropixels Data Analysis

    davila7/claude-code-templates

    Analyzes Neuropixels recordings from SpikeGLX or Open Ephys through preprocessing, drift correction, Kilosort4 spike sorting, quality metrics and curation.

    32k GitHub starsUsed in 10 repos~2.8k tokens
    Auto-check passed
  • Scientific Venue Templates

    davila7/claude-code-templates

    Supplies LaTeX templates and formatting rules for journals, conferences, posters, and grant proposals, then can check a draft against them.

    32k GitHub starsUsed in 9 repos~5.1k tokens
    Auto-check: notes
  • Brand Voice Content Creator

    davila7/claude-code-templates

    Analyzes a brand's existing writing to lock in a consistent voice, then builds SEO blog posts and platform-specific social content around it.

    32k GitHub starsUsed in 3 repos~1.9k tokens
    Auto-check passed
  • CAPA Officer

    davila7/claude-code-templates

    Guides corrective and preventive action (CAPA) work in a quality management system, from initiation and root cause analysis through effectiveness verification.

    32k GitHub starsUsed in 1 repo~2k tokens
    Auto-check passed
  • Fda Consultant Specialist

    davila7/claude-code-templates

    Senior FDA consultant and specialist for medical device companies including HIPAA compliance and requirement management.

    32k GitHub starsUsed in 1 repo~2.7k tokens
    Auto-check passed

Questions about PyTorch Lightning Training Setup

What does PyTorch Lightning Training Setup do?

Organizes PyTorch training code into LightningModules, DataModules and Trainers, with multi-GPU strategies, callbacks and logging configured. PyTorch Lightning cuts training boilerplate while keeping control of the model. The skill shows the agent how to split a model into a LightningModule's six sections (initialization, training, validation, test and prediction steps, and optimizer configuration), and how to wrap data handling in a LightningDataModule with prepare_data, setup and the train, validation and test dataloaders.

When should I use PyTorch Lightning Training Setup?

PyTorch Lightning Training Setup fits situations like: restructuring a plain PyTorch training loop into a LightningModule; configuring a Trainer for multi-GPU or TPU training; putting dataset loading and splits into a reusable DataModule; adding checkpointing, early stopping and logging to a training run.

How do I install PyTorch Lightning Training Setup in Claude Code?

Run `npx skills add davila7/claude-code-templates --skill pytorch-lightning -a claude-code`. Or copy the skill folder (cli-tool/components/skills/scientific/pytorch-lightning in davila7/claude-code-templates) into .claude/skills/pytorch-lightning in your project. Claude Code loads it when a task matches its description.

How do I install PyTorch Lightning Training Setup in Codex?

Run `npx skills add davila7/claude-code-templates --skill pytorch-lightning -a codex`. Or copy the skill folder (cli-tool/components/skills/scientific/pytorch-lightning in davila7/claude-code-templates) into .agents/skills/pytorch-lightning in your project. Codex loads it when a task matches its description.

Can I use PyTorch Lightning Training Setup in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add davila7/claude-code-templates --skill pytorch-lightning -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/pytorch-lightning, .gemini/skills/pytorch-lightning, .github/skills/pytorch-lightning and .opencode/skills/pytorch-lightning in your project.

What does PyTorch Lightning Training Setup need to run?

Going by SKILL.md and its folder, PyTorch Lightning Training Setup needs Python for the scripts in its folder. Our summary lists: PyTorch and PyTorch Lightning installed; Multiple GPUs or TPUs for distributed runs (optional).

Does PyTorch Lightning Training Setup access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is PyTorch Lightning Training Setup safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does PyTorch Lightning Training Setup use?

PyTorch Lightning Training Setup is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does PyTorch Lightning Training Setup use?

About 1.7k tokens (SKILL.md is roughly 6.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 26k tokens, read only when the agent opens those files.

What are the alternatives to PyTorch Lightning Training Setup?

Skills that share tags, products or a category with PyTorch Lightning Training Setup: PyTorch Lightning Training (Orchestra-Research/AI-Research-SKILLs, 13k stars), Ray Train Distributed Training (Orchestra-Research/AI-Research-SKILLs, 13k stars), GPU Optimizer (Mathews-Tom/armory, 328 stars) and Cuda Index Width (pytorch/pytorch, 104k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains PyTorch Lightning Training Setup?

davila7 (a GitHub user) maintains it in davila7/claude-code-templates, which has 32,463 GitHub stars. The repository holds 477 skills in this directory. The repository was last updated on October 8, 2026.

Source: davila7/claude-code-templates on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.