Agent skill

Torchtitan

by Luciole-Studio in Luciole-Studio/Misaka-Agent

Pretrain LLMs at scale with PyTorch 4D parallelism. An agent skill from Luciole-Studio/Misaka-Agent.

MITAuto-check passedAI & LLM Engineering

Install Torchtitan

skills CLI
$ npx skills add Luciole-Studio/Misaka-Agent --skill torchtitan -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Luciole-Studio/Misaka-Agent torchtitan --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Luciole-Studio/Misaka-Agent.git skills-src && mkdir -p .claude/skills && cp -r skills-src/misaka/core/skills/assets/optional/mlops/torchtitan .claude/skills/torchtitan && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
torchtitan
GitHub stars
158
Used in
1 other repo
Token cost
~2.6k tokens
SKILL.md length
543 words
Files
5 (incl. references)
Skills in repo
77
Repo updated
First seen
Licence
MIT

At a glance

Pretrain LLMs at scale with PyTorch 4D parallelism. An agent skill from Luciole-Studio/Misaka-Agent.

  • Tasks that involve Deep learning
  • SKILL.md covers Quick start, Common workflows, When to use vs alternatives and Common issues, plus 4 more sections
  • Calls pip, python and git; reaches github.com and huggingface.co
  • Tasks that involve MLOps

What it does

Torchtitan is an agent skill from Luciole-Studio/Misaka-Agent. Pretrain LLMs at scale with PyTorch 4D parallelism.

Its SKILL.md is about 2.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including reference files (for example `references/checkpoint.md`, `references/custom-models.md` and `references/float8.md`).

It sits in AI & LLM Engineering, covering Deep learning and MLOps. It works with PyTorch. The repository describes itself as: A multi-agent research system for the humanities and social sciences. The licence is MIT.

When your agent uses it

  • Tasks that involve Deep learning
  • Tasks that involve MLOps

Example prompts

  • “/torchtitan”

Requirements

  • Python 3
  • A credential in YOUR_HF_TOKEN

What it can do on your machine

Read from SKILL.md and the folder at commit 3bcf7a3. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pip
    • python
    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • github.com
    • huggingface.co

    Also links to:

    • arxiv.org
    • iclr.cc
    • discuss.pytorch.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Torchtitan loads about 2.6k tokens when it runs, and up to ~7.5k if it reads all its reference files. Until then it costs about 16 tokens; SKILL.md has 543 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~16
When it runs · the whole SKILL.md, loaded when a task matches
~2.6k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~7.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Luciole-Studio/Misaka-Agent at commit 3bcf7a3, republished under its MIT licence (© Luciole-Studio). 543 words, ~2,584 tokens.

Download SKILL.mdSave it as .claude/skills/torchtitan/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
torchtitan
description
Pretrain LLMs at scale with PyTorch 4D parallelism.
version
1.0.1
author
Orchestra Research
license
MIT
dependencies
torch>=2.6.0, torchtitan>=0.2.0, torchao>=0.5.0
platforms
linux, macos

TorchTitan - PyTorch Native Distributed LLM Pretraining

Quick start

TorchTitan is PyTorch's official platform for large-scale LLM pretraining with composable 4D parallelism (FSDP2, TP, PP, CP), achieving 65%+ speedups over baselines on H100 GPUs.

Installation:

bash
# From PyPI (stable)
pip install torchtitan

# From source (latest features, requires PyTorch nightly)
git clone https://github.com/pytorch/torchtitan
cd torchtitan
pip install -r requirements.txt

Download tokenizer:

bash
# Get HF token from https://huggingface.co/settings/tokens
python scripts/download_hf_assets.py --repo_id meta-llama/Llama-3.1-8B --assets tokenizer --hf_token=...

Start training on 8 GPUs:

bash
# Configs are selected by name from the Python config registry
# (torchtitan/models/llama3/config_registry.py), not by TOML path
MODULE=llama3 CONFIG=llama3_8b ./run_train.sh

Common workflows

Workflow 1: Pretrain Llama 3.1 8B on single node

Copy this checklist:

Single Node Pretraining:
- [ ] Step 1: Download tokenizer
- [ ] Step 2: Configure training
- [ ] Step 3: Launch training
- [ ] Step 4: Monitor and checkpoint

Step 1: Download tokenizer

bash
python scripts/download_hf_assets.py \
  --repo_id meta-llama/Llama-3.1-8B \
  --assets tokenizer \
  --hf_token=YOUR_HF_TOKEN

Step 2: Configure training

In torchtitan's current layout, run configs are defined in a Python config registry (torchtitan/models/llama3/config_registry.py) and selected by name via CONFIG=<name> (or --config <name>). To customize, register your own config in the registry, or override individual fields on the command line (e.g. --optimizer.lr 3e-4 --training.steps 1000).

The equivalent settings for an 8B run look like this (shown as fields; set them in the registry entry or as --section.key value overrides):

toml
# fields for a llama3 8B run (register in config_registry.py or pass as --overrides)
[job]
dump_folder = "./outputs"
description = "Llama 3.1 8B training"

[model]
name = "llama3"
flavor = "8B"
hf_assets_path = "./assets/hf/Llama-3.1-8B"

[optimizer]
name = "AdamW"
lr = 3e-4

[lr_scheduler]
warmup_steps = 200

[training]
local_batch_size = 2
seq_len = 8192
max_norm = 1.0
steps = 1000
dataset = "c4"

[parallelism]
data_parallel_shard_degree = -1  # Use all GPUs for FSDP

[activation_checkpoint]
mode = "selective"
selective_ac_option = "op"

[checkpoint]
enable = true
folder = "checkpoint"
interval = 500

Step 3: Launch training

bash
# 8 GPUs on single node (config selected by name from the registry)
MODULE=llama3 CONFIG=llama3_8b ./run_train.sh

# Override individual fields on the command line
MODULE=llama3 CONFIG=llama3_8b ./run_train.sh --optimizer.lr 3e-4 --training.steps 1000

# Or explicitly with torchrun (run_train.sh wraps this)
torchrun --nproc_per_node=8 \
  -m torchtitan.train \
  --module llama3 --config llama3_8b

Step 4: Monitor and checkpoint

TensorBoard logs are saved to ./outputs/tb/:

bash
tensorboard --logdir ./outputs/tb
Workflow 2: Multi-node training with SLURM
Multi-Node Training:
- [ ] Step 1: Configure parallelism for scale
- [ ] Step 2: Set up SLURM script
- [ ] Step 3: Submit job
- [ ] Step 4: Resume from checkpoint

Step 1: Configure parallelism for scale

For 70B model on 256 GPUs (32 nodes):

toml
[parallelism]
data_parallel_shard_degree = 32  # FSDP across 32 ranks
tensor_parallel_degree = 8        # TP within node
pipeline_parallel_degree = 1      # No PP for 70B
context_parallel_degree = 1       # Increase for long sequences

Step 2: Set up SLURM script

bash
#!/bin/bash
#SBATCH --job-name=llama70b
#SBATCH --nodes=32
#SBATCH --ntasks-per-node=8
#SBATCH --gpus-per-node=8

srun torchrun \
  --nnodes=32 \
  --nproc_per_node=8 \
  --rdzv_backend=c10d \
  --rdzv_endpoint=$MASTER_ADDR:$MASTER_PORT \
  -m torchtitan.train \
  --module llama3 --config llama3_70b

Step 3: Submit job

bash
sbatch multinode_trainer.slurm

Step 4: Resume from checkpoint

Training auto-resumes if checkpoint exists in configured folder.

Workflow 3: Enable Float8 training for H100s

Float8 provides 30-50% speedup on H100 GPUs.

Float8 Training:
- [ ] Step 1: Install torchao
- [ ] Step 2: Configure Float8
- [ ] Step 3: Launch with compile

Step 1: Install torchao

bash
USE_CPP=0 pip install git+https://github.com/pytorch/ao.git

Step 2: Configure Float8

In the current torchtitan, Float8 is applied at config time via the quantization parameter in your model_registry() call inside the config registry (not via a [quantize.linear.float8] TOML section). Add a Float8LinearConverter.Config:

python
# in torchtitan/models/llama3/config_registry.py (your model_registry(...) call)
from torchtitan.components.quantization import Float8LinearConverter

model_spec = model_registry(
    "8B",
    quantization=[
        Float8LinearConverter.Config(
            recipe_name="rowwise",          # or "rowwise_with_gw_hp"
            filter_fqns=["output"],          # skip layers too small to benefit
            model_compile_enabled=True,      # requires torch.compile for competitive perf
        ),
    ],
)

Enable torch.compile in your run config too:

toml
[compile]
enable = true
components = ["model", "loss"]

Step 3: Launch with compile

bash
# Float8 config is baked into the registered config; just select it and enable compile
MODULE=llama3 CONFIG=llama3_8b ./run_train.sh --compile.enable
Workflow 4: 4D parallelism for 405B models
4D Parallelism (FSDP + TP + PP + CP):
- [ ] Step 1: Create seed checkpoint
- [ ] Step 2: Configure 4D parallelism
- [ ] Step 3: Launch on 512 GPUs

Step 1: Create seed checkpoint

Required for consistent initialization across PP stages:

bash
NGPU=1 MODULE=llama3 CONFIG=llama3_405b ./run_train.sh \
  --checkpoint.enable \
  --checkpoint.create_seed_checkpoint \
  --parallelism.data_parallel_shard_degree 1 \
  --parallelism.tensor_parallel_degree 1 \
  --parallelism.pipeline_parallel_degree 1

Step 2: Configure 4D parallelism

toml
[parallelism]
data_parallel_shard_degree = 8   # FSDP
tensor_parallel_degree = 8       # TP within node
pipeline_parallel_degree = 8     # PP across nodes
context_parallel_degree = 1      # CP for long sequences

[training]
local_batch_size = 32
seq_len = 8192

Step 3: Launch on 512 GPUs

bash
# 64 nodes x 8 GPUs = 512 GPUs
srun torchrun --nnodes=64 --nproc_per_node=8 \
  -m torchtitan.train \
  --module llama3 --config llama3_405b

When to use vs alternatives

Use TorchTitan when:

  • Pretraining LLMs from scratch (8B to 405B+)
  • Need PyTorch-native solution without third-party dependencies
  • Require composable 4D parallelism (FSDP2, TP, PP, CP)
  • Training on H100s with Float8 support
  • Want interoperable checkpoints with torchtune/HuggingFace

Use alternatives instead:

  • Megatron-LM: Maximum performance for NVIDIA-only deployments
  • DeepSpeed: Broader ZeRO optimization ecosystem, inference support
  • Axolotl/TRL: Fine-tuning rather than pretraining
  • LitGPT: Educational, smaller-scale training
Show full SKILL.md (193 more words)Show less

Common issues

Issue: Out of memory on large models

Enable activation checkpointing and reduce batch size:

toml
[activation_checkpoint]
mode = "full"  # Instead of "selective"

[training]
local_batch_size = 1

Or use gradient accumulation:

toml
[training]
local_batch_size = 1
global_batch_size = 32  # Accumulates gradients

Issue: TP causes high memory with async collectives

Set environment variable:

bash
export TORCH_NCCL_AVOID_RECORD_STREAMS=1

Issue: Float8 training not faster

Float8 only benefits large GEMMs. Filter small layers via the converter's filter_fqns:

python
from torchtitan.components.quantization import Float8LinearConverter

Float8LinearConverter.Config(
    # add "auto_filter_small_kn" to auto-skip layers too small to benefit
    filter_fqns=["attention.wk", "attention.wv", "output", "auto_filter_small_kn"],
    model_compile_enabled=True,
)

Issue: Checkpoint loading fails after parallelism change

Use DCP's resharding capability:

bash
# Convert sharded checkpoint to single file
python -m torch.distributed.checkpoint.format_utils \
  dcp_to_torch checkpoint/step-1000 checkpoint.pt

Issue: Pipeline parallelism initialization

Create seed checkpoint first (see Workflow 4, Step 1).

Supported models

ModelSizesStatus
Llama 3.18B, 70B, 405BProduction
Llama 4VariousExperimental
DeepSeek V316B, 236B, 671B (MoE)Experimental
GPT-OSS20B, 120B (MoE)Experimental
Qwen 3VariousExperimental
FluxDiffusionExperimental

Performance benchmarks (H100)

ModelGPUsParallelismTPS/GPUTechniques
Llama 8B8FSDP5,762Baseline
Llama 8B8FSDP+compile+FP88,532+48%
Llama 70B256FSDP+TP+AsyncTP8762D parallel
Llama 405B512FSDP+TP+PP1283D parallel

Advanced topics

FSDP2 configuration: See references/fsdp.md for detailed FSDP2 vs FSDP1 comparison and ZeRO equivalents.

Float8 training: See references/float8.md for tensorwise vs rowwise scaling recipes.

Checkpointing: See references/checkpoint.md for HuggingFace conversion and async checkpointing.

Adding custom models: See references/custom-models.md for TrainSpec protocol.

Resources

© Luciole-Studio, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (references) in misaka/core/skills/assets/optional/mlops/torchtitan of Luciole-Studio/Misaka-Agent.

  • SKILL.md
  • references/checkpoint.md
  • references/custom-models.md
  • references/float8.md
  • references/fsdp.md

Open the folder on GitHubat commit 3bcf7a3

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in Luciole-Studio/Misaka-Agent, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Torchtitan next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Torchtitan compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Torchtitan this skillLuciole-Studio/Misaka-Agent1581 repos~2.6kAutomated safety check: PassMIT
TensorBoard Training VisualizationOrchestra-Research/AI-Research-SKILLs13k3 repos~3.8kAutomated safety check: PassMIT
Senior ML Engineerdavila7/claude-code-templates32k2 repos~1.4kAutomated safety check: PassMIT
Databricks ML Trainingdatabricks/databricks-agent-skills345—~4.6kAutomated safety check: PassCustom licence
ML EngineerRightNow-AI/openfang18k—~987Automated safety check: PassApache-2.0
ML Experimentrevfactory/harness-1001.3k—~1.9kAutomated safety check: PassApache-2.0

Similar skills

  • TensorBoard Training Visualization

    Orchestra-Research/AI-Research-SKILLs

    Covers logging and viewing training metrics, histograms, model graphs, embeddings and profiles with TensorBoard in PyTorch and TensorFlow projects.

    13k GitHub starsUsed in 3 repos~3.8k tokens
    AI & LLM EngineeringAuto-check passed
  • Senior ML Engineer

    davila7/claude-code-templates

    World-class ML engineering skill for productionizing ML models, MLOps, and building scalable ML systems.

    32k GitHub starsUsed in 2 repos~1.4k tokens
    AI & LLM EngineeringAuto-check passed
  • Databricks ML Training

    databricks/databricks-agent-skills

    Official

    Train ML models on Databricks. An agent skill from databricks/databricks-agent-skills.

    345 GitHub stars~4.6k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • ML Engineer

    RightNow-AI/openfang

    Machine learning engineer expert for PyTorch, scikit-learn, model evaluation, and MLOps

    18k GitHub stars~987 tokensUpdated 3 mo ago
    Data & AnalyticsAuto-check passed
  • ML Experiment

    revfactory/harness-100

    A full ML pipeline where an agent team collaborates to perform data preparation, model design, training, evaluation, and deployment readiness.

    1.3k GitHub stars~1.9k tokensUpdated 6 mo ago
    Data & AnalyticsAuto-check passed
  • Edit

    omegaml/omegaml

    how to use the edit command properly

    107 GitHub stars~206 tokensUpdated today
    DevOps & CloudAuto-check passed

More from Luciole-Studio/Misaka-Agent

All 77 skills in this repo
  • Kanban Video Orchestrator

    Luciole-Studio/Misaka-Agent

    Plan and run multi-agent video production pipelines. An agent skill from Luciole-Studio/Misaka-Agent.

    158 GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check: notes
  • Ast Grep

    Luciole-Studio/Misaka-Agent

    AST-aware structural code search and rewrite via ast-grep. An agent skill from Luciole-Studio/Misaka-Agent.

    158 GitHub starsUsed in 1 repo~3.2k tokens
    Auto-check passed
  • Drug Discovery

    Luciole-Studio/Misaka-Agent

    Drug discovery: ChEMBL search, drug-likeness, interactions. An agent skill from Luciole-Studio/Misaka-Agent.

    158 GitHub starsUsed in 1 repo~2.2k tokens
    Auto-check passed
  • Fitness Nutrition

    Luciole-Studio/Misaka-Agent

    Workout planning, macros, and body metrics via wger/USDA. An agent skill from Luciole-Studio/Misaka-Agent.

    158 GitHub starsUsed in 1 repo~2.4k tokens
    Auto-check passed
  • Hyperframes

    Luciole-Studio/Misaka-Agent

    Render MP4/WebM videos from HTML compositions. An agent skill from Luciole-Studio/Misaka-Agent.

    158 GitHub starsUsed in 1 repo~3.9k tokens
    Auto-check passed
  • Osint Investigation

    Luciole-Studio/Misaka-Agent

    Follow the money via public records and sanctions data. An agent skill from Luciole-Studio/Misaka-Agent.

    158 GitHub starsUsed in 1 repo~2.9k tokens
    Auto-check passed

Works with

Questions about Torchtitan

What does Torchtitan do?

Pretrain LLMs at scale with PyTorch 4D parallelism. An agent skill from Luciole-Studio/Misaka-Agent. Torchtitan is an agent skill from Luciole-Studio/Misaka-Agent. Pretrain LLMs at scale with PyTorch 4D parallelism.

When should I use Torchtitan?

Torchtitan fits situations like: tasks that involve Deep learning; tasks that involve MLOps.

How do I install Torchtitan in Claude Code?

Run `npx skills add Luciole-Studio/Misaka-Agent --skill torchtitan -a claude-code`. Or copy the skill folder (misaka/core/skills/assets/optional/mlops/torchtitan in Luciole-Studio/Misaka-Agent) into .claude/skills/torchtitan in your project. Claude Code loads it when a task matches its description.

How do I install Torchtitan in Codex?

Run `npx skills add Luciole-Studio/Misaka-Agent --skill torchtitan -a codex`. Or copy the skill folder (misaka/core/skills/assets/optional/mlops/torchtitan in Luciole-Studio/Misaka-Agent) into .agents/skills/torchtitan in your project. Codex loads it when a task matches its description.

Can I use Torchtitan in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Luciole-Studio/Misaka-Agent --skill torchtitan -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/torchtitan, .gemini/skills/torchtitan, .github/skills/torchtitan and .opencode/skills/torchtitan in your project.

What does Torchtitan need to run?

Going by SKILL.md and its folder, Torchtitan needs the command-line tools its instructions call (pip, python and git). Our summary lists: Python 3; A credential in YOUR_HF_TOKEN.

Does Torchtitan access the network?

SKILL.md names 5 domains. In commands or code: github.com and huggingface.co; the agent is likely to contact these when it follows the instructions. As links in the text: arxiv.org, iclr.cc and discuss.pytorch.org. This is read from the text; nothing was executed.

Is Torchtitan safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Torchtitan use?

Torchtitan is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Torchtitan use?

About 2.6k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4.9k tokens, read only when the agent opens those files.

What are the alternatives to Torchtitan?

Skills that share tags, products or a category with Torchtitan: TensorBoard Training Visualization (Orchestra-Research/AI-Research-SKILLs, 13k stars), Senior ML Engineer (davila7/claude-code-templates, 32k stars), Databricks ML Training (databricks/databricks-agent-skills, 345 stars) and ML Engineer (RightNow-AI/openfang, 18k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Torchtitan?

Luciole-Studio (a GitHub organization) maintains it in Luciole-Studio/Misaka-Agent, which has 158 GitHub stars. The repository holds 77 skills in this directory. The repository was last updated on October 8, 2026.

Source: Luciole-Studio/Misaka-Agent on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.