Docstring
pytorch/pytorch
Write docstrings for PyTorch functions and methods following PyTorch conventions.
Agent skill
by Orchestra-Research in Orchestra-Research/AI-Research-SKILLs
Provides PyTorch-native distributed LLM pretraining using torchtitan with 4D parallelism (FSDP2, TP, PP, CP).
$ npx skills add Orchestra-Research/AI-Research-SKILLs --skill distributed-llm-pretraining-torchtitan -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install Orchestra-Research/AI-Research-SKILLs distributed-llm-pretraining-torchtitan --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/Orchestra-Research/AI-Research-SKILLs.git skills-src && mkdir -p .claude/skills && cp -r skills-src/01-model-architecture/torchtitan .claude/skills/distributed-llm-pretraining-torchtitan && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "distributed-llm-pretraining-torchtitan" agent skill from https://github.com/Orchestra-Research/AI-Research-SKILLs/tree/main/01-model-architecture/torchtitan into .claude/skills/distributed-llm-pretraining-torchtitan/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "distributed-llm-pretraining-torchtitan", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/Orchestra-Research/AI-Research-SKILLs/tree/main/01-model-architecture/torchtitanType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add Orchestra-Research/AI-Research-SKILLs --skill distributed-llm-pretraining-torchtitan -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install Orchestra-Research/AI-Research-SKILLs distributed-llm-pretraining-torchtitan --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Orchestra-Research/AI-Research-SKILLs.git skills-src && mkdir -p .agents/skills && cp -r skills-src/01-model-architecture/torchtitan .agents/skills/distributed-llm-pretraining-torchtitan && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "distributed-llm-pretraining-torchtitan" agent skill from https://github.com/Orchestra-Research/AI-Research-SKILLs/tree/main/01-model-architecture/torchtitan into .agents/skills/distributed-llm-pretraining-torchtitan/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "distributed-llm-pretraining-torchtitan", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Orchestra-Research/AI-Research-SKILLs --skill distributed-llm-pretraining-torchtitan -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install Orchestra-Research/AI-Research-SKILLs distributed-llm-pretraining-torchtitan --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Orchestra-Research/AI-Research-SKILLs.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/01-model-architecture/torchtitan .cursor/skills/distributed-llm-pretraining-torchtitan && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "distributed-llm-pretraining-torchtitan" agent skill from https://github.com/Orchestra-Research/AI-Research-SKILLs/tree/main/01-model-architecture/torchtitan into .cursor/skills/distributed-llm-pretraining-torchtitan/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "distributed-llm-pretraining-torchtitan", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/Orchestra-Research/AI-Research-SKILLs.git --path 01-model-architecture/torchtitan--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add Orchestra-Research/AI-Research-SKILLs --skill distributed-llm-pretraining-torchtitan -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install Orchestra-Research/AI-Research-SKILLs distributed-llm-pretraining-torchtitan --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Orchestra-Research/AI-Research-SKILLs.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/01-model-architecture/torchtitan .gemini/skills/distributed-llm-pretraining-torchtitan && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "distributed-llm-pretraining-torchtitan" agent skill from https://github.com/Orchestra-Research/AI-Research-SKILLs/tree/main/01-model-architecture/torchtitan into .gemini/skills/distributed-llm-pretraining-torchtitan/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "distributed-llm-pretraining-torchtitan", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install Orchestra-Research/AI-Research-SKILLs distributed-llm-pretraining-torchtitanInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add Orchestra-Research/AI-Research-SKILLs --skill distributed-llm-pretraining-torchtitan -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/Orchestra-Research/AI-Research-SKILLs.git skills-src && mkdir -p .github/skills && cp -r skills-src/01-model-architecture/torchtitan .github/skills/distributed-llm-pretraining-torchtitan && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "distributed-llm-pretraining-torchtitan" agent skill from https://github.com/Orchestra-Research/AI-Research-SKILLs/tree/main/01-model-architecture/torchtitan into .github/skills/distributed-llm-pretraining-torchtitan/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "distributed-llm-pretraining-torchtitan", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Orchestra-Research/AI-Research-SKILLs --skill distributed-llm-pretraining-torchtitan -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install Orchestra-Research/AI-Research-SKILLs distributed-llm-pretraining-torchtitan --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Orchestra-Research/AI-Research-SKILLs.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/01-model-architecture/torchtitan .opencode/skills/distributed-llm-pretraining-torchtitan && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "distributed-llm-pretraining-torchtitan" agent skill from https://github.com/Orchestra-Research/AI-Research-SKILLs/tree/main/01-model-architecture/torchtitan into .opencode/skills/distributed-llm-pretraining-torchtitan/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "distributed-llm-pretraining-torchtitan", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
distributed-llm-pretraining-torchtitanProvides PyTorch-native distributed LLM pretraining using torchtitan with 4D parallelism (FSDP2, TP, PP, CP).
Distributed LLM Pretraining Torchtitan is an agent skill from Orchestra-Research/AI-Research-SKILLs. Provides PyTorch-native distributed LLM pretraining using torchtitan with 4D parallelism (FSDP2, TP, PP, CP). Use when pretraining Llama 3.1, DeepSeek V3, or custom models at scale from 8 to 512+ GPUs with Float8, torch.compile, and distributed checkpointing.
Its SKILL.md is about 2.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including reference files (for example `references/checkpoint.md`, `references/custom-models.md` and `references/float8.md`).
It sits in AI & LLM Engineering, covering Deep learning. It works with PyTorch and DeepSeek. The repository describes itself as: Comprehensive open-source library of AI research and engineering skills for any AI model. Package the skills and your claude code/codex/gemini agent will be an AI research agent… The licence is MIT.
Read from SKILL.md and the folder at commit 773a529. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pippythongitFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
github.comhuggingface.coAlso links to:
arxiv.orgiclr.ccdiscuss.pytorch.orgFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Distributed LLM Pretraining Torchtitan loads about 2.2k tokens when it runs, and up to ~7.1k if it reads all its reference files. Until then it costs about 75 tokens; SKILL.md has 444 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from Orchestra-Research/AI-Research-SKILLs at commit 773a529, republished under its MIT licence (© Orchestra-Research). 444 words, ~2,232 tokens.
.claude/skills/distributed-llm-pretraining-torchtitan/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.TorchTitan is PyTorch's official platform for large-scale LLM pretraining with composable 4D parallelism (FSDP2, TP, PP, CP), achieving 65%+ speedups over baselines on H100 GPUs.
Installation:
# From PyPI (stable)
pip install torchtitan
# From source (latest features, requires PyTorch nightly)
git clone https://github.com/pytorch/torchtitan
cd torchtitan
pip install -r requirements.txtDownload tokenizer:
# Get HF token from https://huggingface.co/settings/tokens
python scripts/download_hf_assets.py --repo_id meta-llama/Llama-3.1-8B --assets tokenizer --hf_token=...Start training on 8 GPUs:
CONFIG_FILE="./torchtitan/models/llama3/train_configs/llama3_8b.toml" ./run_train.shCopy this checklist:
Single Node Pretraining:
- [ ] Step 1: Download tokenizer
- [ ] Step 2: Configure training
- [ ] Step 3: Launch training
- [ ] Step 4: Monitor and checkpointStep 1: Download tokenizer
python scripts/download_hf_assets.py \
--repo_id meta-llama/Llama-3.1-8B \
--assets tokenizer \
--hf_token=YOUR_HF_TOKENStep 2: Configure training
Edit or create a TOML config file:
# llama3_8b_custom.toml
[job]
dump_folder = "./outputs"
description = "Llama 3.1 8B training"
[model]
name = "llama3"
flavor = "8B"
hf_assets_path = "./assets/hf/Llama-3.1-8B"
[optimizer]
name = "AdamW"
lr = 3e-4
[lr_scheduler]
warmup_steps = 200
[training]
local_batch_size = 2
seq_len = 8192
max_norm = 1.0
steps = 1000
dataset = "c4"
[parallelism]
data_parallel_shard_degree = -1 # Use all GPUs for FSDP
[activation_checkpoint]
mode = "selective"
selective_ac_option = "op"
[checkpoint]
enable = true
folder = "checkpoint"
interval = 500Step 3: Launch training
# 8 GPUs on single node
CONFIG_FILE="./llama3_8b_custom.toml" ./run_train.sh
# Or explicitly with torchrun
torchrun --nproc_per_node=8 \
-m torchtitan.train \
--job.config_file ./llama3_8b_custom.tomlStep 4: Monitor and checkpoint
TensorBoard logs are saved to ./outputs/tb/:
tensorboard --logdir ./outputs/tbMulti-Node Training:
- [ ] Step 1: Configure parallelism for scale
- [ ] Step 2: Set up SLURM script
- [ ] Step 3: Submit job
- [ ] Step 4: Resume from checkpointStep 1: Configure parallelism for scale
For 70B model on 256 GPUs (32 nodes):
[parallelism]
data_parallel_shard_degree = 32 # FSDP across 32 ranks
tensor_parallel_degree = 8 # TP within node
pipeline_parallel_degree = 1 # No PP for 70B
context_parallel_degree = 1 # Increase for long sequencesStep 2: Set up SLURM script
#!/bin/bash
#SBATCH --job-name=llama70b
#SBATCH --nodes=32
#SBATCH --ntasks-per-node=8
#SBATCH --gpus-per-node=8
srun torchrun \
--nnodes=32 \
--nproc_per_node=8 \
--rdzv_backend=c10d \
--rdzv_endpoint=$MASTER_ADDR:$MASTER_PORT \
-m torchtitan.train \
--job.config_file ./llama3_70b.tomlStep 3: Submit job
sbatch multinode_trainer.slurmStep 4: Resume from checkpoint
Training auto-resumes if checkpoint exists in configured folder.
Float8 provides 30-50% speedup on H100 GPUs.
Float8 Training:
- [ ] Step 1: Install torchao
- [ ] Step 2: Configure Float8
- [ ] Step 3: Launch with compileStep 1: Install torchao
USE_CPP=0 pip install git+https://github.com/pytorch/ao.gitStep 2: Configure Float8
Add to your TOML config:
[model]
converters = ["quantize.linear.float8"]
[quantize.linear.float8]
enable_fsdp_float8_all_gather = true
precompute_float8_dynamic_scale_for_fsdp = true
filter_fqns = ["output"] # Exclude output layer
[compile]
enable = true
components = ["model", "loss"]Step 3: Launch with compile
CONFIG_FILE="./llama3_8b.toml" ./run_train.sh \
--model.converters="quantize.linear.float8" \
--quantize.linear.float8.enable_fsdp_float8_all_gather \
--compile.enable4D Parallelism (FSDP + TP + PP + CP):
- [ ] Step 1: Create seed checkpoint
- [ ] Step 2: Configure 4D parallelism
- [ ] Step 3: Launch on 512 GPUsStep 1: Create seed checkpoint
Required for consistent initialization across PP stages:
NGPU=1 CONFIG_FILE=./llama3_405b.toml ./run_train.sh \
--checkpoint.enable \
--checkpoint.create_seed_checkpoint \
--parallelism.data_parallel_shard_degree 1 \
--parallelism.tensor_parallel_degree 1 \
--parallelism.pipeline_parallel_degree 1Step 2: Configure 4D parallelism
[parallelism]
data_parallel_shard_degree = 8 # FSDP
tensor_parallel_degree = 8 # TP within node
pipeline_parallel_degree = 8 # PP across nodes
context_parallel_degree = 1 # CP for long sequences
[training]
local_batch_size = 32
seq_len = 8192Step 3: Launch on 512 GPUs
# 64 nodes x 8 GPUs = 512 GPUs
srun torchrun --nnodes=64 --nproc_per_node=8 \
-m torchtitan.train \
--job.config_file ./llama3_405b.tomlUse TorchTitan when:
Use alternatives instead:
Issue: Out of memory on large models
Enable activation checkpointing and reduce batch size:
[activation_checkpoint]
mode = "full" # Instead of "selective"
[training]
local_batch_size = 1Or use gradient accumulation:
[training]
local_batch_size = 1
global_batch_size = 32 # Accumulates gradientsIssue: TP causes high memory with async collectives
Set environment variable:
export TORCH_NCCL_AVOID_RECORD_STREAMS=1Issue: Float8 training not faster
Float8 only benefits large GEMMs. Filter small layers:
[quantize.linear.float8]
filter_fqns = ["attention.wk", "attention.wv", "output", "auto_filter_small_kn"]Issue: Checkpoint loading fails after parallelism change
Use DCP's resharding capability:
# Convert sharded checkpoint to single file
python -m torch.distributed.checkpoint.format_utils \
dcp_to_torch checkpoint/step-1000 checkpoint.ptIssue: Pipeline parallelism initialization
Create seed checkpoint first (see Workflow 4, Step 1).
| Model | Sizes | Status |
|---|---|---|
| Llama 3.1 | 8B, 70B, 405B | Production |
| Llama 4 | Various | Experimental |
| DeepSeek V3 | 16B, 236B, 671B (MoE) | Experimental |
| GPT-OSS | 20B, 120B (MoE) | Experimental |
| Qwen 3 | Various | Experimental |
| Flux | Diffusion | Experimental |
| Model | GPUs | Parallelism | TPS/GPU | Techniques |
|---|---|---|---|---|
| Llama 8B | 8 | FSDP | 5,762 | Baseline |
| Llama 8B | 8 | FSDP+compile+FP8 | 8,532 | +48% |
| Llama 70B | 256 | FSDP+TP+AsyncTP | 876 | 2D parallel |
| Llama 405B | 512 | FSDP+TP+PP | 128 | 3D parallel |
FSDP2 configuration: See references/fsdp.md for detailed FSDP2 vs FSDP1 comparison and ZeRO equivalents.
Float8 training: See references/float8.md for tensorwise vs rowwise scaling recipes.
Checkpointing: See references/checkpoint.md for HuggingFace conversion and async checkpointing.
Adding custom models: See references/custom-models.md for TrainSpec protocol.
© Orchestra-Research, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 4 other files (references) in 01-model-architecture/torchtitan of Orchestra-Research/AI-Research-SKILLs.
Open the folder on GitHubat commit 773a529
We found 4 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 4 other GitHub owners. This page covers the copy in Orchestra-Research/AI-Research-SKILLs, which our catalogue first saw on October 7, 2026.
Distributed LLM Pretraining Torchtitan next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Distributed LLM Pretraining Torchtitan this skillOrchestra-Research/AI-Research-SKILLs | 13k | 4 repos | ~2.2k | Automated safety check: Pass | MIT | |
| Docstringpytorch/pytorch | 104k | 2 repos | ~2.6k | Automated safety check: Pass | Custom licence | |
| CI Metricspytorch/pytorch | 104k | — | ~1.1k | Automated safety check: Pass | Custom licence | |
| Cuda Index Widthpytorch/pytorch | 104k | — | ~1.6k | Automated safety check: Pass | Custom licence | |
| Benchmark Pyreflyfacebook/pyrefly | 7.1k | — | ~1.8k | Automated safety check: Pass | MIT | |
| Document Public APIspytorch/pytorch | 104k | — | ~4.2k | Automated safety check: Pass | Custom licence |
pytorch/pytorch
Write docstrings for PyTorch functions and methods following PyTorch conventions.
pytorch/pytorch
Query PyTorch CI, GitHub Actions, HUD, Grafana, and infrastructure metrics.
pytorch/pytorch
Choose 32-bit vs 64-bit index math in PyTorch CUDA kernels. An agent skill from pytorch/pytorch.
facebook/pyrefly
Run Pyrefly benchmarks locally via Buck or Cargo, including PyTorch real-world LSP benchmarks.
pytorch/pytorch
Document undocumented public APIs in PyTorch by removing functions from coverageignorefunctions and coverageignoreclasses in docs/source/conf.py, running Sphinx coverage, and adding the appropriate…
linkedin/Liger-Kernel
Develops production-ready Triton kernels for Liger Kernel. An agent skill from linkedin/Liger-Kernel.
Orchestra-Research/AI-Research-SKILLs
Generates music from text descriptions with MusicGen and sound effects with AudioGen, using Meta's AudioCraft PyTorch library with melody and style conditioning.
Orchestra-Research/AI-Research-SKILLs
Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.
Orchestra-Research/AI-Research-SKILLs
Guide to using Meta's Segment Anything Model for zero-shot image segmentation with point, box or mask prompts, or automatic mask generation.
Orchestra-Research/AI-Research-SKILLs
Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.
Orchestra-Research/AI-Research-SKILLs
Explains OpenAI's CLIP model for zero-shot image classification, image-text similarity, semantic image search and content moderation, with install steps and code patterns.
Orchestra-Research/AI-Research-SKILLs
Transcribes audio with OpenAI's Whisper: 99 languages, translation to English, language detection, six model sizes and word-level timestamps, from Python or the CLI.
Categories
Provides PyTorch-native distributed LLM pretraining using torchtitan with 4D parallelism (FSDP2, TP, PP, CP). Distributed LLM Pretraining Torchtitan is an agent skill from Orchestra-Research/AI-Research-SKILLs. Provides PyTorch-native distributed LLM pretraining using torchtitan with 4D parallelism (FSDP2, TP, PP, CP).
Distributed LLM Pretraining Torchtitan fits situations like: pretraining Llama 3.1; custom models at scale from 8 to 512+ GPUs with Float8; distributed checkpointing.
Run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill distributed-llm-pretraining-torchtitan -a claude-code`. Or copy the skill folder (01-model-architecture/torchtitan in Orchestra-Research/AI-Research-SKILLs) into .claude/skills/distributed-llm-pretraining-torchtitan in your project. Claude Code loads it when a task matches its description.
Run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill distributed-llm-pretraining-torchtitan -a codex`. Or copy the skill folder (01-model-architecture/torchtitan in Orchestra-Research/AI-Research-SKILLs) into .agents/skills/distributed-llm-pretraining-torchtitan in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill distributed-llm-pretraining-torchtitan -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/distributed-llm-pretraining-torchtitan, .gemini/skills/distributed-llm-pretraining-torchtitan, .github/skills/distributed-llm-pretraining-torchtitan and .opencode/skills/distributed-llm-pretraining-torchtitan in your project.
Going by SKILL.md and its folder, Distributed LLM Pretraining Torchtitan needs the command-line tools its instructions call (pip, python and git). Our summary lists: Python 3; A credential in YOUR_HF_TOKEN.
SKILL.md names 5 domains. In commands or code: github.com and huggingface.co; the agent is likely to contact these when it follows the instructions. As links in the text: arxiv.org, iclr.cc and discuss.pytorch.org. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Distributed LLM Pretraining Torchtitan is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.2k tokens (SKILL.md is roughly 8.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4.9k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Distributed LLM Pretraining Torchtitan: Docstring (pytorch/pytorch, 104k stars), CI Metrics (pytorch/pytorch, 104k stars), Cuda Index Width (pytorch/pytorch, 104k stars) and Benchmark Pyrefly (facebook/pyrefly, 7.1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Orchestra-Research (a GitHub organization) maintains it in Orchestra-Research/AI-Research-SKILLs, which has 13,405 GitHub stars. The repository holds 96 skills in this directory. The repository was last updated on June 16, 2026.
Source: Orchestra-Research/AI-Research-SKILLs on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.