Adds PyTorch FSDP2 (fullyshard) to training scripts with correct init, sharding, mixed precision/offload config, and distributed checkpointing.

MITAuto-check passedAI & LLM Engineering

Install Pytorch Fsdp2

skills CLI
$ npx skills add Orchestra-Research/AI-Research-SKILLs --skill pytorch-fsdp2 -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Orchestra-Research/AI-Research-SKILLs pytorch-fsdp2 --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Orchestra-Research/AI-Research-SKILLs.git skills-src && mkdir -p .claude/skills && cp -r skills-src/08-distributed-training/pytorch-fsdp2 .claude/skills/pytorch-fsdp2 && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
pytorch-fsdp2
GitHub stars
13k
Used in
1 other repo
Token cost
~2.7k tokens
SKILL.md length
1,078 words
Files
13 (incl. references)
Skills in repo
96
Repo updated
First seen
Licence
MIT

At a glance

Adds PyTorch FSDP2 (fullyshard) to training scripts with correct init, sharding, mixed precision/offload config, and distributed checkpointing.

  • Works in 8 steps: Version & environment sanity → Initialize distributed and set device → Build model on meta device (recommended… → …
  • Models exceed single-GPU memory
  • SKILL.md covers When to use this skill, Alternatives (when FSDP2 is…, Contract the agent must follow and Step-by-step procedure, plus 5 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Pytorch Fsdp2 is an agent skill from Orchestra-Research/AI-Research-SKILLs. Adds PyTorch FSDP2 (fullyshard) to training scripts with correct init, sharding, mixed precision/offload config, and distributed checkpointing. Use when models exceed single-GPU memory or when you need DTensor-based sharding with DeviceMesh.

Its SKILL.md is about 2.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 13 other files, including reference files (for example `references/pytorch_dcp_async_recipe.md`, `references/pytorch_dcp_overview.md` and `references/pytorch_dcp_recipe.md`).

It sits in AI & LLM Engineering, covering Deep learning and Database administration. It works with PyTorch. The repository describes itself as: Comprehensive open-source library of AI research and engineering skills for any AI model. Package the skills and your claude code/codex/gemini agent will be an AI research agent… The licence is MIT.

When your agent uses it

  • Models exceed single-GPU memory
  • You need DTensor-based sharding with DeviceMesh

Example prompts

  • “/pytorch-fsdp2”

Workflow steps

8 steps, taken from the step headings in SKILL.md.

  1. Version & environment sanity
  2. Initialize distributed and set device
  3. Build model on meta device (recommended for very large models)
  4. Apply fully_shard() bottom-up (wrapping policy = “apply where needed”)
  5. Configure reshard_after_forward for memory/perf trade-offs
  6. Mixed precision & offload (optional but common)
  7. Optimizer, gradient clipping, accumulation
  8. Checkpointing: prefer DCP or distributed state dict helpers

What it can do on your machine

Read from SKILL.md and the folder at commit 773a529. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Pytorch Fsdp2 loads about 2.7k tokens when it runs, and up to ~6.1k if it reads all its reference files. Until then it costs about 64 tokens; SKILL.md has 1,078 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~64
When it runs · the whole SKILL.md, loaded when a task matches
~2.7k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~6.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Orchestra-Research/AI-Research-SKILLs at commit 773a529, republished under its MIT licence (© Orchestra-Research). 1,078 words, ~2,737 tokens.

Download SKILL.mdSave it as .claude/skills/pytorch-fsdp2/SKILL.md (or your agent's skills folder). This skill also uses 12 other files; get the full folder from GitHub.
name
pytorch-fsdp2
description
Adds PyTorch FSDP2 (fully_shard) to training scripts with correct init, sharding, mixed precision/offload config, and distributed checkpointing. Use when models exceed single-GPU memory or when you need DTensor-based sharding with DeviceMesh.
version
1.0.0
author
Orchestra Research
license
MIT
tags
PyTorch, FSDP2, Fully Sharded Data Parallel, Distributed Training, DTensor, Device Mesh, Sharded Checkpointing, Mixed Precision, Offload, Torch Distributed
dependencies
torch

Skill: Use PyTorch FSDP2 (fully_shard) correctly in a training script

This skill teaches a coding agent how to add PyTorch FSDP2 to a training loop with correct initialization, sharding, mixed precision/offload configuration, and checkpointing.

FSDP2 in PyTorch is exposed primarily via torch.distributed.fsdp.fully_shard and the FSDPModule methods it adds in-place to modules. See: references/pytorch_fully_shard_api.md, references/pytorch_fsdp2_tutorial.md.


When to use this skill

Use FSDP2 when:

  • Your model doesn’t fit on one GPU (parameters + gradients + optimizer state).
  • You want an eager-mode sharding approach that is DTensor-based per-parameter sharding (more inspectable, simpler sharded state dicts) than FSDP1.
  • You may later compose DP with Tensor Parallel using DeviceMesh.

Avoid (or be careful) if:

  • You need strict backwards-compatible checkpoints across PyTorch versions (DCP warns against this).
  • You’re forced onto older PyTorch versions without the FSDP2 stack.

Alternatives (when FSDP2 is not the best fit)

  • DistributedDataParallel (DDP): Use the standard data-parallel wrapper when you want classic distributed data parallel training.
  • FullyShardedDataParallel (FSDP1): Use the original FSDP wrapper for parameter sharding across data-parallel workers.

Reference: references/pytorch_ddp_notes.md, references/pytorch_fsdp1_api.md.


Contract the agent must follow

  1. Launch with torchrun and set the CUDA device per process (usually via LOCAL_RANK).
  2. Apply fully_shard() bottom-up, i.e., shard submodules (e.g., Transformer blocks) before the root module.
  3. Call model(input), not model.forward(input), so the FSDP2 hooks run (unless you explicitly unshard() or register the forward method).
  4. Create the optimizer after sharding and make sure it is built on the DTensor parameters (post-fully_shard).
  5. Checkpoint using Distributed Checkpoint (DCP) or the distributed-state-dict helpers, not naïve torch.save(model.state_dict()) unless you deliberately gather to full tensors.

(Each of these rules is directly described in the official API docs/tutorial; see references.)


Step-by-step procedure

0) Version & environment sanity
  • Prefer a recent stable PyTorch where the docs show FSDP2 and DCP updated recently.
  • Use torchrun --nproc_per_node <gpus_per_node> ... and ensure RANK, WORLD_SIZE, LOCAL_RANK are visible.

Reference: references/pytorch_fsdp2_tutorial.md (launch commands and setup), references/pytorch_fully_shard_api.md (user contract).


1) Initialize distributed and set device

Minimal, correct pattern:

  • dist.init_process_group(backend="nccl")
  • torch.cuda.set_device(int(os.environ["LOCAL_RANK"]))
  • Optionally create a DeviceMesh to describe the data-parallel group(s)

Reference: references/pytorch_device_mesh_tutorial.md (why DeviceMesh exists & how it manages process groups).


For big models, initialize on meta, apply sharding, then materialize weights on GPU:

  • with torch.device("meta"): model = ...
  • apply fully_shard(...) on submodules, then fully_shard(model)
  • model.to_empty(device="cuda")
  • model.reset_parameters() (or your init routine)

Reference: references/pytorch_fsdp2_tutorial.md (migration guide shows this flow explicitly).


3) Apply fully_shard() bottom-up (wrapping policy = “apply where needed”)

Do not only call fully_shard on the topmost module.

Recommended sharding pattern for transformer-like models:

  • iterate modules, if isinstance(m, TransformerBlock): fully_shard(m, ...)
  • then fully_shard(model, ...)

Why:

  • fully_shard forms “parameter groups” for collective efficiency and excludes params already grouped by earlier calls. Bottom-up gives better overlap and lower peak memory.

Reference: references/pytorch_fully_shard_api.md (bottom-up requirement and why).


4) Configure reshard_after_forward for memory/perf trade-offs

Default behavior:

  • None means True for non-root modules and False for root modules (good default).

Heuristics:

  • If you’re memory-bound: keep defaults or force True on many blocks.
  • If you’re throughput-bound and can afford memory: consider keeping unsharded params longer (root often False).
  • Advanced: use an int to reshard to a smaller mesh after forward (e.g., intra-node) if it’s a meaningful divisor.

Reference: references/pytorch_fully_shard_api.md (full semantics).


5) Mixed precision & offload (optional but common)

FSDP2 uses:

  • mp_policy=MixedPrecisionPolicy(param_dtype=..., reduce_dtype=..., output_dtype=..., cast_forward_inputs=...)
  • offload_policy=CPUOffloadPolicy() if you want CPU offload

Rules of thumb:

  • Start with BF16 parameters/reductions on H100/A100-class GPUs (if numerically stable for your model).
  • Keep reduce_dtype aligned with your gradient reduction expectations.
  • If you use CPU offload, budget for PCIe/NVLink traffic and runtime overhead.

Reference: references/pytorch_fully_shard_api.md (MixedPrecisionPolicy / OffloadPolicy classes).


6) Optimizer, gradient clipping, accumulation
  • Create the optimizer after sharding so it holds DTensor params.
  • If you need gradient accumulation / no_sync:
    • use the FSDP2 mechanism (set_requires_gradient_sync) instead of FSDP1’s no_sync().

Gradient clipping:

  • Use the approach shown in the FSDP2 tutorial (“Gradient Clipping and Optimizer with DTensor”), because parameters/gradients are DTensors.

Reference: references/pytorch_fsdp2_tutorial.md.


Show full SKILL.md (445 more words)Show less
7) Checkpointing: prefer DCP or distributed state dict helpers

Two recommended approaches:

A) Distributed Checkpoint (DCP) — best default

  • DCP saves/loads from multiple ranks in parallel and supports load-time resharding.
  • DCP produces multiple files (often at least one per rank) and operates “in place”.

B) Distributed state dict helpers

  • get_model_state_dict / set_model_state_dict with StateDictOptions(full_state_dict=True, cpu_offload=True, broadcast_from_rank0=True, ...)
  • For optimizer: get_optimizer_state_dict / set_optimizer_state_dict

Avoid:

  • Saving DTensor state dicts with plain torch.save unless you intentionally convert with DTensor.full_tensor() and manage memory carefully.

References:

  • references/pytorch_dcp_overview.md (DCP behavior and caveats)
  • references/pytorch_dcp_recipe.md and references/pytorch_dcp_async_recipe.md (end-to-end usage)
  • references/pytorch_fsdp2_tutorial.md (DTensor vs DCP state-dict flows)
  • references/pytorch_examples_fsdp2.md (working checkpoint scripts)

Workflow checklists (copy-paste friendly)

Workflow A: Retrofit FSDP2 into an existing training script
  • Launch with torchrun and initialize the process group.
  • Set the CUDA device from LOCAL_RANK; create a DeviceMesh if you need multi-dim parallelism.
  • Build the model (use meta if needed), apply fully_shard bottom-up, then fully_shard(model).
  • Create the optimizer after sharding so it captures DTensor parameters.
  • Use model(inputs) so hooks run; use set_requires_gradient_sync for accumulation.
  • Add DCP save/load via torch.distributed.checkpoint helpers.

Reference: references/pytorch_fsdp2_tutorial.md, references/pytorch_fully_shard_api.md, references/pytorch_device_mesh_tutorial.md, references/pytorch_dcp_recipe.md.

Workflow B: Add DCP save/load (minimal pattern)
  • Wrap state in Stateful or assemble state via get_state_dict.
  • Call dcp.save(...) from all ranks to a shared path.
  • Call dcp.load(...) and restore with set_state_dict.
  • Validate any resharding assumptions when loading into a different mesh.

Reference: references/pytorch_dcp_recipe.md.

Debug checklist (what the agent should check first)

  1. All ranks on distinct GPUs?
    If not, verify torch.cuda.set_device(LOCAL_RANK) and your torchrun flags.
  2. Did you accidentally call forward() directly?
    Use model(input) or explicitly unshard() / register forward.
  3. Is fully_shard() applied bottom-up?
    If only root is sharded, expect worse memory/perf and possible confusion.
  4. Optimizer created at the right time?
    Must be built on DTensor parameters after sharding.
  5. Checkpointing path consistent?
    • If using DCP, don’t mix with ad-hoc torch.save unless you understand conversions.
    • Be mindful of PyTorch-version compatibility warnings for DCP.

Common issues and fixes

  • Forward hooks not running → Call model(inputs) (or unshard() explicitly) instead of model.forward(...).
  • Optimizer sees non-DTensor params → Create optimizer after all fully_shard calls.
  • Only root module sharded → Apply fully_shard bottom-up on submodules before the root.
  • Memory spikes after forward → Set reshard_after_forward=True for more modules.
  • Gradient accumulation desync → Use set_requires_gradient_sync instead of FSDP1’s no_sync().

Reference: references/pytorch_fully_shard_api.md, references/pytorch_fsdp2_tutorial.md.


Minimal reference implementation outline (agent-friendly)

The coding agent should implement a script with these labeled blocks:

  • init_distributed(): init process group, set device
  • build_model_meta(): model on meta, apply fully_shard, materialize weights
  • build_optimizer(): optimizer created after sharding
  • train_step(): forward/backward/step with model(inputs) and DTensor-aware patterns
  • checkpoint_save/load(): DCP or distributed state dict helpers

Concrete examples live in references/pytorch_examples_fsdp2.md and the official tutorial reference.


References

  • references/pytorch_fsdp2_tutorial.md
  • references/pytorch_fully_shard_api.md
  • references/pytorch_ddp_notes.md
  • references/pytorch_fsdp1_api.md
  • references/pytorch_device_mesh_tutorial.md
  • references/pytorch_tp_tutorial.md
  • references/pytorch_dcp_overview.md
  • references/pytorch_dcp_recipe.md
  • references/pytorch_dcp_async_recipe.md
  • references/pytorch_examples_fsdp2.md
  • references/torchtitan_fsdp_notes.md (optional, production notes)
  • references/ray_train_fsdp2_example.md (optional, integration example)

© Orchestra-Research, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 12 other files (references) in 08-distributed-training/pytorch-fsdp2 of Orchestra-Research/AI-Research-SKILLs.

  • SKILL.md
  • references/pytorch_dcp_async_recipe.md
  • references/pytorch_dcp_overview.md
  • references/pytorch_dcp_recipe.md
  • references/pytorch_ddp_notes.md
  • references/pytorch_device_mesh_tutorial.md
  • references/pytorch_examples_fsdp2.md
  • references/pytorch_fsdp1_api.md
  • references/pytorch_fsdp2_tutorial.md
  • references/pytorch_fully_shard_api.md
  • references/pytorch_tp_tutorial.md
  • references/ray_train_fsdp2_example.md
  • references/torchtitan_fsdp_notes.md

Open the folder on GitHubat commit 773a529

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in Orchestra-Research/AI-Research-SKILLs, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Pytorch Fsdp2 next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Pytorch Fsdp2 compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Pytorch Fsdp2 this skillOrchestra-Research/AI-Research-SKILLs13k1 repos~2.7kAutomated safety check: PassMIT
Physicsnemo Shard TensorNVIDIA/skills3.5k—~3.5kAutomated safety check: PassApache-2.0
Add Uint Supportpytorch/pytorch104k2 repos~2.3kAutomated safety check: PassCustom licence
Add Torch Shapes Examplefacebook/pyrefly7.1k—~1.3kAutomated safety check: PassMIT
MUSA GPU Training Optimizeropen-infra-skills/infra-skills141—~1.7kAutomated safety check: PassApache-2.0
Ghstack CIpytorch/pytorch104k—~1.4kAutomated safety check: PassCustom licence

Similar skills

  • Official

    Official NVIDIA-authored guidance for PhysicsNeMo ShardTensor domain parallelism — integrate domain parallelism into training/inference scripts (new or existing) with DDP or FSDP2, write and…

    3.5k GitHub stars~3.5k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Add Uint Support

    pytorch/pytorch

    Add unsigned integer (uint) type support to PyTorch operators by updating ATDISPATCH macros.

    104k GitHub starsUsed in 2 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Add Torch Shapes Example

    facebook/pyrefly

    Official

    A skill your agent uses when adding a new PyTorch model to Pyrefly's shape-tracking example corpus under tensor-shapes/pyrefly-torch-stubs/examples — i.e.

    7.1k GitHub stars~1.3k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • MUSA GPU Training Optimizer

    open-infra-skills/infra-skills

    Profiles, benchmarks and tunes AI training workloads on Moore Threads MUSA GPUs with a measurement-first process that keeps model behavior unchanged.

    141 GitHub stars~1.7k tokensUpdated 3 mo ago
    AI & LLM EngineeringAuto-check passed
  • Ghstack CI

    pytorch/pytorch

    Manage CI for PyTorch ghstack stacks by running CI where its results are useful now and deferring other PRs with [no-ci].

    104k GitHub stars~1.4k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Depth Estimation

    SharpAI/DeepCamera

    Real-time depth map privacy transforms using Depth Anything v2 (CoreML + PyTorch)

    3.1k GitHub stars~945 tokensUpdated 23 days ago
    AI & LLM EngineeringAuto-check passed

More from Orchestra-Research/AI-Research-SKILLs

All 96 skills in this repo
  • AudioCraft Audio Generation

    Orchestra-Research/AI-Research-SKILLs

    Generates music from text descriptions with MusicGen and sound effects with AudioGen, using Meta's AudioCraft PyTorch library with melody and style conditioning.

    13k GitHub starsUsed in 8 repos~3.9k tokens
    Auto-check passed
  • LLM Benchmarking with lm-evaluation-harness

    Orchestra-Research/AI-Research-SKILLs

    Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.

    13k GitHub starsUsed in 8 repos~3k tokens
    Auto-check passed
  • Segment Anything Model Guide

    Orchestra-Research/AI-Research-SKILLs

    Guide to using Meta's Segment Anything Model for zero-shot image segmentation with point, box or mask prompts, or automatic mask generation.

    13k GitHub starsUsed in 8 repos~3.3k tokens
    Auto-check passed
  • Chroma Vector Database

    Orchestra-Research/AI-Research-SKILLs

    Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.

    13k GitHub starsUsed in 7 repos~2.3k tokens
    Auto-check passed
  • CLIP Image-Text Matching

    Orchestra-Research/AI-Research-SKILLs

    Explains OpenAI's CLIP model for zero-shot image classification, image-text similarity, semantic image search and content moderation, with install steps and code patterns.

    13k GitHub starsUsed in 7 repos~1.7k tokens
    Auto-check passed
  • Whisper Speech Recognition

    Orchestra-Research/AI-Research-SKILLs

    Transcribes audio with OpenAI's Whisper: 99 languages, translation to English, language detection, six model sizes and word-level timestamps, from Python or the CLI.

    13k GitHub starsUsed in 7 repos~1.9k tokens
    Auto-check: notes

Works with

Questions about Pytorch Fsdp2

What does Pytorch Fsdp2 do?

Adds PyTorch FSDP2 (fullyshard) to training scripts with correct init, sharding, mixed precision/offload config, and distributed checkpointing. Pytorch Fsdp2 is an agent skill from Orchestra-Research/AI-Research-SKILLs. Adds PyTorch FSDP2 (fullyshard) to training scripts with correct init, sharding, mixed precision/offload config, and distributed checkpointing.

When should I use Pytorch Fsdp2?

Pytorch Fsdp2 fits situations like: models exceed single-GPU memory; you need DTensor-based sharding with DeviceMesh.

How do I install Pytorch Fsdp2 in Claude Code?

Run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill pytorch-fsdp2 -a claude-code`. Or copy the skill folder (08-distributed-training/pytorch-fsdp2 in Orchestra-Research/AI-Research-SKILLs) into .claude/skills/pytorch-fsdp2 in your project. Claude Code loads it when a task matches its description.

How do I install Pytorch Fsdp2 in Codex?

Run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill pytorch-fsdp2 -a codex`. Or copy the skill folder (08-distributed-training/pytorch-fsdp2 in Orchestra-Research/AI-Research-SKILLs) into .agents/skills/pytorch-fsdp2 in your project. Codex loads it when a task matches its description.

Can I use Pytorch Fsdp2 in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill pytorch-fsdp2 -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/pytorch-fsdp2, .gemini/skills/pytorch-fsdp2, .github/skills/pytorch-fsdp2 and .opencode/skills/pytorch-fsdp2 in your project.

What does Pytorch Fsdp2 need to run?

SKILL.md names no scripts, command-line tools or credentials: Pytorch Fsdp2 is instructions for the agent only.

Does Pytorch Fsdp2 access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Pytorch Fsdp2 safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Pytorch Fsdp2 use?

Pytorch Fsdp2 is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Pytorch Fsdp2 use?

About 2.7k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.3k tokens, read only when the agent opens those files.

What are the alternatives to Pytorch Fsdp2?

Skills that share tags, products or a category with Pytorch Fsdp2: Physicsnemo Shard Tensor (NVIDIA/skills, 3.5k stars), Add Uint Support (pytorch/pytorch, 104k stars), Add Torch Shapes Example (facebook/pyrefly, 7.1k stars) and MUSA GPU Training Optimizer (open-infra-skills/infra-skills, 141 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Pytorch Fsdp2?

Orchestra-Research (a GitHub organization) maintains it in Orchestra-Research/AI-Research-SKILLs, which has 13,374 GitHub stars. The repository holds 96 skills in this directory. The repository was last updated on June 16, 2026.

Source: Orchestra-Research/AI-Research-SKILLs on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.