Agent skill

Hugging Face Accelerate

by Orchestra-Research in Orchestra-Research/AI-Research-SKILLs

Adds distributed and mixed-precision training to a PyTorch script with a few Accelerate lines, then launches it on one GPU, many GPUs or DeepSpeed and FSDP setups.

MITAuto-check passedAI & LLM Engineering

Install Hugging Face Accelerate

skills CLI
$ npx skills add Orchestra-Research/AI-Research-SKILLs --skill huggingface-accelerate -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Orchestra-Research/AI-Research-SKILLs huggingface-accelerate --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Orchestra-Research/AI-Research-SKILLs.git skills-src && mkdir -p .claude/skills && cp -r skills-src/08-distributed-training/accelerate .claude/skills/huggingface-accelerate && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
huggingface-accelerate
GitHub stars
13k
Used in
5 other repos
Token cost
~2.1k tokens
SKILL.md length
328 words
Files
4 (incl. references)
Skills in repo
96
Repo updated
First seen
Licence
MIT

At a glance

Adds distributed and mixed-precision training to a PyTorch script with a few Accelerate lines, then launches it on one GPU, many GPUs or DeepSpeed and FSDP setups.

  • Scaling a single-GPU PyTorch script to several GPUs
  • SKILL.md covers Quick start, Common workflows, When to use vs alternatives and Common issues, plus 3 more sections
  • Calls pip
  • Turning on mixed precision or gradient accumulation in a training loop

What it does

The skill centers on converting an existing PyTorch training script by importing Accelerator and wrapping the model, optimizer and data loaders, a change of about four lines. An interactive accelerate config asks about hardware, machine count, mixed precision and DeepSpeed, and accelerate launch then runs the same script on a single GPU, several GPUs or multiple machines without further code changes.

Worked workflows cover moving from one GPU to several, enabling FP16 or BF16 mixed precision, DeepSpeed ZeRO-2 through code or a JSON config file, FSDP with a plugin object, and gradient accumulation, with the effective batch size given as batch size times GPU count times accumulation steps. A section compares when Accelerate fits and when alternatives are better. Reference files cover custom plugins, Megatron integration and performance.

When your agent uses it

  • Scaling a single-GPU PyTorch script to several GPUs
  • Turning on mixed precision or gradient accumulation in a training loop
  • Switching a training script between DDP, DeepSpeed and FSDP
  • Writing one script that runs on any hardware

Example prompts

  • “Convert train.py to use Accelerate so it runs on multiple GPUs with bf16 mixed precision.”
  • “Set up DeepSpeed ZeRO-2 for my training script through an accelerate config file.”
  • “Add gradient accumulation to this loop and tell me the effective batch size on four GPUs.”

Requirements

  • Python with PyTorch and the accelerate package
  • GPUs, TPUs or CPUs to run the launch command on

What it can do on your machine

Read from SKILL.md and the folder at commit 773a529. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com
    • huggingface.co

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Hugging Face Accelerate loads about 2.1k tokens when it runs, and up to ~11k if it reads all its reference files. Until then it costs about 74 tokens; SKILL.md has 328 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~74
When it runs · the whole SKILL.md, loaded when a task matches
~2.1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~11k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Orchestra-Research/AI-Research-SKILLs at commit 773a529, republished under its MIT licence (© Orchestra-Research). 328 words, ~2,084 tokens.

Download SKILL.mdSave it as .claude/skills/huggingface-accelerate/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
huggingface-accelerate
description
Simplest distributed training API. 4 lines to add distributed support to any PyTorch script. Unified API for DeepSpeed/FSDP/Megatron/DDP. Automatic device placement, mixed precision (FP16/BF16/FP8). Interactive config, single launch command. HuggingFace ecosystem standard.
version
1.0.0
author
Orchestra Research
license
MIT
tags
Distributed Training, HuggingFace, Accelerate, DeepSpeed, FSDP, Mixed Precision, PyTorch, DDP, Unified API, Simple
dependencies
accelerate, torch, transformers

HuggingFace Accelerate - Unified Distributed Training

Quick start

Accelerate simplifies distributed training to 4 lines of code.

Installation:

bash
pip install accelerate

Convert PyTorch script (4 lines):

python
import torch
+ from accelerate import Accelerator

+ accelerator = Accelerator()

  model = torch.nn.Transformer()
  optimizer = torch.optim.Adam(model.parameters())
  dataloader = torch.utils.data.DataLoader(dataset)

+ model, optimizer, dataloader = accelerator.prepare(model, optimizer, dataloader)

  for batch in dataloader:
      optimizer.zero_grad()
      loss = model(batch)
-     loss.backward()
+     accelerator.backward(loss)
      optimizer.step()

Run (single command):

bash
accelerate launch train.py

Common workflows

Workflow 1: From single GPU to multi-GPU

Original script:

python
# train.py
import torch

model = torch.nn.Linear(10, 2).to('cuda')
optimizer = torch.optim.Adam(model.parameters())
dataloader = torch.utils.data.DataLoader(dataset, batch_size=32)

for epoch in range(10):
    for batch in dataloader:
        batch = batch.to('cuda')
        optimizer.zero_grad()
        loss = model(batch).mean()
        loss.backward()
        optimizer.step()

With Accelerate (4 lines added):

python
# train.py
import torch
from accelerate import Accelerator  # +1

accelerator = Accelerator()  # +2

model = torch.nn.Linear(10, 2)
optimizer = torch.optim.Adam(model.parameters())
dataloader = torch.utils.data.DataLoader(dataset, batch_size=32)

model, optimizer, dataloader = accelerator.prepare(model, optimizer, dataloader)  # +3

for epoch in range(10):
    for batch in dataloader:
        # No .to('cuda') needed - automatic!
        optimizer.zero_grad()
        loss = model(batch).mean()
        accelerator.backward(loss)  # +4
        optimizer.step()

Configure (interactive):

bash
accelerate config

Questions:

  • Which machine? (single/multi GPU/TPU/CPU)
  • How many machines? (1)
  • Mixed precision? (no/fp16/bf16/fp8)
  • DeepSpeed? (no/yes)

Launch (works on any setup):

bash
# Single GPU
accelerate launch train.py

# Multi-GPU (8 GPUs)
accelerate launch --multi_gpu --num_processes 8 train.py

# Multi-node
accelerate launch --multi_gpu --num_processes 16 \
  --num_machines 2 --machine_rank 0 \
  --main_process_ip $MASTER_ADDR \
  train.py
Workflow 2: Mixed precision training

Enable FP16/BF16:

python
from accelerate import Accelerator

# FP16 (with gradient scaling)
accelerator = Accelerator(mixed_precision='fp16')

# BF16 (no scaling, more stable)
accelerator = Accelerator(mixed_precision='bf16')

# FP8 (H100+)
accelerator = Accelerator(mixed_precision='fp8')

model, optimizer, dataloader = accelerator.prepare(model, optimizer, dataloader)

# Everything else is automatic!
for batch in dataloader:
    with accelerator.autocast():  # Optional, done automatically
        loss = model(batch)
    accelerator.backward(loss)
Workflow 3: DeepSpeed ZeRO integration

Enable DeepSpeed ZeRO-2:

python
from accelerate import Accelerator

accelerator = Accelerator(
    mixed_precision='bf16',
    deepspeed_plugin={
        "zero_stage": 2,  # ZeRO-2
        "offload_optimizer": False,
        "gradient_accumulation_steps": 4
    }
)

# Same code as before!
model, optimizer, dataloader = accelerator.prepare(model, optimizer, dataloader)

Or via config:

bash
accelerate config
# Select: DeepSpeed → ZeRO-2

deepspeed_config.json:

json
{
    "fp16": {"enabled": false},
    "bf16": {"enabled": true},
    "zero_optimization": {
        "stage": 2,
        "offload_optimizer": {"device": "cpu"},
        "allgather_bucket_size": 5e8,
        "reduce_bucket_size": 5e8
    }
}

Launch:

bash
accelerate launch --config_file deepspeed_config.json train.py
Workflow 4: FSDP (Fully Sharded Data Parallel)

Enable FSDP:

python
from accelerate import Accelerator, FullyShardedDataParallelPlugin

fsdp_plugin = FullyShardedDataParallelPlugin(
    sharding_strategy="FULL_SHARD",  # ZeRO-3 equivalent
    auto_wrap_policy="TRANSFORMER_AUTO_WRAP",
    cpu_offload=False
)

accelerator = Accelerator(
    mixed_precision='bf16',
    fsdp_plugin=fsdp_plugin
)

model, optimizer, dataloader = accelerator.prepare(model, optimizer, dataloader)

Or via config:

bash
accelerate config
# Select: FSDP → Full Shard → No CPU Offload
Workflow 5: Gradient accumulation

Accumulate gradients:

python
from accelerate import Accelerator

accelerator = Accelerator(gradient_accumulation_steps=4)

model, optimizer, dataloader = accelerator.prepare(model, optimizer, dataloader)

for batch in dataloader:
    with accelerator.accumulate(model):  # Handles accumulation
        optimizer.zero_grad()
        loss = model(batch)
        accelerator.backward(loss)
        optimizer.step()

Effective batch size: batch_size * num_gpus * gradient_accumulation_steps

When to use vs alternatives

Use Accelerate when:

  • Want simplest distributed training
  • Need single script for any hardware
  • Use HuggingFace ecosystem
  • Want flexibility (DDP/DeepSpeed/FSDP/Megatron)
  • Need quick prototyping

Key advantages:

  • 4 lines: Minimal code changes
  • Unified API: Same code for DDP, DeepSpeed, FSDP, Megatron
  • Automatic: Device placement, mixed precision, sharding
  • Interactive config: No manual launcher setup
  • Single launch: Works everywhere

Use alternatives instead:

  • PyTorch Lightning: Need callbacks, high-level abstractions
  • Ray Train: Multi-node orchestration, hyperparameter tuning
  • DeepSpeed: Direct API control, advanced features
  • Raw DDP: Maximum control, minimal abstraction

Common issues

Issue: Wrong device placement

Don't manually move to device:

python
# WRONG
batch = batch.to('cuda')

# CORRECT
# Accelerate handles it automatically after prepare()

Issue: Gradient accumulation not working

Use context manager:

python
# CORRECT
with accelerator.accumulate(model):
    optimizer.zero_grad()
    accelerator.backward(loss)
    optimizer.step()

Issue: Checkpointing in distributed

Use accelerator methods:

python
# Save only on main process
if accelerator.is_main_process:
    accelerator.save_state('checkpoint/')

# Load on all processes
accelerator.load_state('checkpoint/')

Issue: Different results with FSDP

Ensure same random seed:

python
from accelerate.utils import set_seed
set_seed(42)

Advanced topics

Megatron integration: See references/megatron-integration.md for tensor parallelism, pipeline parallelism, and sequence parallelism setup.

Custom plugins: See references/custom-plugins.md for creating custom distributed plugins and advanced configuration.

Performance tuning: See references/performance.md for profiling, memory optimization, and best practices.

Hardware requirements

  • CPU: Works (slow)
  • Single GPU: Works
  • Multi-GPU: DDP (default), DeepSpeed, or FSDP
  • Multi-node: DDP, DeepSpeed, FSDP, Megatron
  • TPU: Supported
  • Apple MPS: Supported

Launcher requirements:

  • DDP: torch.distributed.run (built-in)
  • DeepSpeed: deepspeed (pip install deepspeed)
  • FSDP: PyTorch 1.12+ (built-in)
  • Megatron: Custom setup

Resources

© Orchestra-Research, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (references) in 08-distributed-training/accelerate of Orchestra-Research/AI-Research-SKILLs.

  • SKILL.md
  • references/custom-plugins.md
  • references/megatron-integration.md
  • references/performance.md

Open the folder on GitHubat commit 773a529

Used in 5 other repositories

We found 5 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 5 other GitHub owners. This page covers the copy in Orchestra-Research/AI-Research-SKILLs, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Hugging Face Accelerate next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Hugging Face Accelerate compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Hugging Face Accelerate this skillOrchestra-Research/AI-Research-SKILLs13k5 repos~2.1kAutomated safety check: PassMIT
Magpie Kernel Evaluatoramd/skills406—~2.3kAutomated safety check: PassMIT
Hyperpod Version Checkerawslabs/agent-plugins915—~910Automated safety check: PassApache-2.0
Triton Langmohitmishra786/low-level-dev-skills253—~1.8kAutomated safety check: PassMIT
MUSA GPU Training Optimizeropen-infra-skills/infra-skills141—~1.7kAutomated safety check: PassApache-2.0
Hugging Face Vision Trainerhuggingface/skills11k1 repos~7.5kAutomated safety check: PassApache-2.0

Similar skills

  • Benchmarks LLM inference and drives GPU kernel optimization with Magpie.

    406 GitHub stars~2.3k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Hyperpod Version Checker

    awslabs/agent-plugins

    Official

    Check and compare software component versions on SageMaker HyperPod cluster nodes - NVIDIA drivers, CUDA toolkit, cuDNN, NCCL, EFA, AWS OFI NCCL, GDRCopy, MPI, Neuron SDK (Trainium/Inferentia)…

    915 GitHub stars~910 tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Triton Lang

    mohitmishra786/low-level-dev-skills

    Triton language skill for Python GPU kernel authoring. An agent skill from mohitmishra786/low-level-dev-skills.

    253 GitHub stars~1.8k tokensUpdated 3 mo ago
    AI & LLM EngineeringAuto-check passed
  • MUSA GPU Training Optimizer

    open-infra-skills/infra-skills

    Profiles, benchmarks and tunes AI training workloads on Moore Threads MUSA GPUs with a measurement-first process that keeps model behavior unchanged.

    141 GitHub stars~1.7k tokensUpdated 3 mo ago
    AI & LLM EngineeringAuto-check passed
  • Hugging Face Vision Trainer

    huggingface/skills

    Official

    Trains and fine-tunes object detection, image classification and SAM or SAM2 segmentation models on Hugging Face Jobs cloud GPUs and saves the results to the Hub.

    11k GitHub starsUsed in 1 repo~7.5k tokens
    AI & LLM EngineeringAuto-check passed
  • Megatron-LM on SLURM

    NVIDIA/Megatron-LM

    Official

    Shows how to launch distributed Megatron-LM training on a SLURM cluster: sbatch skeleton, torch.distributed.run setup, CUDA_DEVICE_MAX_CONNECTIONS rules and failure diagnosis.

    18k GitHub stars~1.8k tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from Orchestra-Research/AI-Research-SKILLs

All 96 skills in this repo
  • AudioCraft Audio Generation

    Orchestra-Research/AI-Research-SKILLs

    Generates music from text descriptions with MusicGen and sound effects with AudioGen, using Meta's AudioCraft PyTorch library with melody and style conditioning.

    13k GitHub starsUsed in 8 repos~3.9k tokens
    Auto-check passed
  • LLM Benchmarking with lm-evaluation-harness

    Orchestra-Research/AI-Research-SKILLs

    Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.

    13k GitHub starsUsed in 8 repos~3k tokens
    Auto-check passed
  • Segment Anything Model Guide

    Orchestra-Research/AI-Research-SKILLs

    Guide to using Meta's Segment Anything Model for zero-shot image segmentation with point, box or mask prompts, or automatic mask generation.

    13k GitHub starsUsed in 8 repos~3.3k tokens
    Auto-check passed
  • Chroma Vector Database

    Orchestra-Research/AI-Research-SKILLs

    Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.

    13k GitHub starsUsed in 7 repos~2.3k tokens
    Auto-check passed
  • CLIP Image-Text Matching

    Orchestra-Research/AI-Research-SKILLs

    Explains OpenAI's CLIP model for zero-shot image classification, image-text similarity, semantic image search and content moderation, with install steps and code patterns.

    13k GitHub starsUsed in 7 repos~1.7k tokens
    Auto-check passed
  • Whisper Speech Recognition

    Orchestra-Research/AI-Research-SKILLs

    Transcribes audio with OpenAI's Whisper: 99 languages, translation to English, language detection, six model sizes and word-level timestamps, from Python or the CLI.

    13k GitHub starsUsed in 7 repos~1.9k tokens
    Auto-check: notes

Questions about Hugging Face Accelerate

What does Hugging Face Accelerate do?

Adds distributed and mixed-precision training to a PyTorch script with a few Accelerate lines, then launches it on one GPU, many GPUs or DeepSpeed and FSDP setups. The skill centers on converting an existing PyTorch training script by importing Accelerator and wrapping the model, optimizer and data loaders, a change of about four lines. An interactive accelerate config asks about hardware, machine count, mixed precision and DeepSpeed, and accelerate launch then runs the same script on a single GPU, several GPUs or multiple machines without further code changes.

When should I use Hugging Face Accelerate?

Hugging Face Accelerate fits situations like: scaling a single-GPU PyTorch script to several GPUs; turning on mixed precision or gradient accumulation in a training loop; switching a training script between DDP, DeepSpeed and FSDP; writing one script that runs on any hardware.

How do I install Hugging Face Accelerate in Claude Code?

Run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill huggingface-accelerate -a claude-code`. Or copy the skill folder (08-distributed-training/accelerate in Orchestra-Research/AI-Research-SKILLs) into .claude/skills/huggingface-accelerate in your project. Claude Code loads it when a task matches its description.

How do I install Hugging Face Accelerate in Codex?

Run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill huggingface-accelerate -a codex`. Or copy the skill folder (08-distributed-training/accelerate in Orchestra-Research/AI-Research-SKILLs) into .agents/skills/huggingface-accelerate in your project. Codex loads it when a task matches its description.

Can I use Hugging Face Accelerate in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill huggingface-accelerate -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/huggingface-accelerate, .gemini/skills/huggingface-accelerate, .github/skills/huggingface-accelerate and .opencode/skills/huggingface-accelerate in your project.

What does Hugging Face Accelerate need to run?

Going by SKILL.md and its folder, Hugging Face Accelerate needs the command-line tools its instructions call (pip). Our summary lists: Python with PyTorch and the accelerate package; GPUs, TPUs or CPUs to run the launch command on.

Does Hugging Face Accelerate access the network?

SKILL.md names 2 domains. As links in the text: github.com and huggingface.co. This is read from the text; nothing was executed.

Is Hugging Face Accelerate safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Hugging Face Accelerate use?

Hugging Face Accelerate is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Hugging Face Accelerate use?

About 2.1k tokens (SKILL.md is roughly 8.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 8.9k tokens, read only when the agent opens those files.

What are the alternatives to Hugging Face Accelerate?

Skills that share tags, products or a category with Hugging Face Accelerate: Magpie Kernel Evaluator (amd/skills, 406 stars), Hyperpod Version Checker (awslabs/agent-plugins, 915 stars), Triton Lang (mohitmishra786/low-level-dev-skills, 253 stars) and MUSA GPU Training Optimizer (open-infra-skills/infra-skills, 141 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Hugging Face Accelerate?

Orchestra-Research (a GitHub organization) maintains it in Orchestra-Research/AI-Research-SKILLs, which has 13,374 GitHub stars. The repository holds 96 skills in this directory. The repository was last updated on June 16, 2026.

Source: Orchestra-Research/AI-Research-SKILLs on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.