Agent skill

Distributed Offline Packing

by EvolvingLMMs-Lab in EvolvingLMMs-Lab/LLaVA-OneVision-2

Bilingual guide for running offlinepacking/autopipe.sh across multiple nodes to produce padding-free packed WebDataset shards for SFT, with Energon Metadataset assembly

Apache-2.0Auto-check passedWriting & Content

Install Distributed Offline Packing

skills CLI
$ npx skills add EvolvingLMMs-Lab/LLaVA-OneVision-2 --skill distributed-offline-packing -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install EvolvingLMMs-Lab/LLaVA-OneVision-2 distributed-offline-packing --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/EvolvingLMMs-Lab/LLaVA-OneVision-2.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.opencode/skills/distributed-offline-packing .claude/skills/distributed-offline-packing && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
distributed-offline-packing
GitHub stars
1.2k
Token cost
~2.6k tokens
SKILL.md length
927 words
Files
1
Skills in repo
8
Repo updated
First seen
Licence
Apache-2.0

At a glance

Bilingual guide for running offlinepacking/autopipe.sh across multiple nodes to produce padding-free packed WebDataset shards for SFT, with Energon Metadataset assembly

  • Works in 7 steps: Pre-flight: choose L / 预检:选 L → Split JSONL across nodes / 切分 JSONL → Mount NFS on every node / 挂载 NFS → …
  • Tasks that involve Translation
  • SKILL.md covers Purpose / 用途, Prerequisites / 前置条件, Architecture / 架构 and Step-by-step Workflow / 操作步骤, plus 4 more sections
  • Calls python and bash

What it does

Distributed Offline Packing is an agent skill from EvolvingLMMs-Lab/LLaVA-OneVision-2. Bilingual guide for running offlinepacking/autopipe.sh across multiple nodes to produce padding-free packed WebDataset shards for SFT, with Energon Metadataset assembly

Its SKILL.md is about 2.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts. Compatibility notes: opencode

It sits in Writing & Content, covering Translation. The repository describes itself as: Fully Open Framework for Democratized Multimodal Training. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Translation

Example prompts

  • “/distributed-offline-packing”

Requirements

  • Python 3
  • Docker
  • Compatibility (from SKILL.md): opencode

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Pre-flight: choose L / 预检:选 L
  2. Split JSONL across nodes / 切分 JSONL
  3. Mount NFS on every node / 挂载 NFS
  4. Launch container on every node / 每台启动容器
  5. Run auto_pipe.sh in parallel on each node / 并行启动
  6. Verify outputs / 验证产物
  7. Write top-level Metadataset yaml / 写顶层 Metadataset yaml

What it can do on your machine

Read from SKILL.md and the folder at commit 6ef16b1. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python
    • bash

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    opencode

    From compatibility in the SKILL.md frontmatter.

Context cost

Distributed Offline Packing loads about 2.6k tokens when it runs. Until then it costs about 50 tokens; SKILL.md has 927 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~50
When it runs · the whole SKILL.md, loaded when a task matches
~2.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from EvolvingLMMs-Lab/LLaVA-OneVision-2 at commit 6ef16b1, republished under its Apache-2.0 licence (© EvolvingLMMs-Lab). 927 words, ~2,570 tokens.

Download SKILL.mdSave it as .claude/skills/distributed-offline-packing/SKILL.md (or your agent's skills folder).
name
distributed-offline-packing
description
Bilingual guide for running offline_packing/auto_pipe.sh across multiple nodes to produce padding-free packed WebDataset shards for SFT, with Energon Metadataset assembly
compatibility
opencode
metadata.domain
data-pipeline
metadata.framework
llava-onevision2
metadata.repo
llava-onevision2

Purpose / 用途

Use this skill when packing a large SFT JSONL (hundreds of thousands to millions of samples) into Energon WebDataset shards at a fixed sequence length, using offline_packing/auto_pipe.sh parallelized across multiple nodes.

当需要把大规模 SFT JSONL(几十万到几百万样本)按固定序列长度打包成 Energon WebDataset shards,并通过 offline_packing/auto_pipe.sh 在多台机器上并行处理时,使用这个 skill。

Prerequisites / 前置条件

  • All nodes share the same NFS mount (data + repo + output)
  • Same docker image on every node (must contain transformers, energon, project repo)
  • offline_packing/auto_pipe.sh and stage scripts s1_split_json_to_samples.py … s4_bins_to_webdataset.py
  • Tokenizer + image processor available locally (HF format model dir)
  • Source JSONL where each line has images, prompts, captions (multi-turn list-of-lists) and image paths are usable as-is

所有节点共享同一个 NFS(数据+代码+输出)。每台机器使用同一个 docker 镜像,里面要有 transformers、energon、项目代码。需要 offline_packing/auto_pipe.sh 和 s1–s4 四个 stage 脚本。Tokenizer 和 image processor 是本地 HF 格式目录。源 JSONL 每行包含 images / prompts / captions(多轮是 list-of-list),图片路径可直接使用。

Architecture / 架构

Pipeline Stages / 流水线阶段
JSONL (N samples)
  ├─ s1_split_json_to_samples.py   # validate + drop bad/missing-image samples
  │                                # output: per-sample serialized records
  ├─ s2_compute_token_lengths.py   # tokenize prompts/captions, compute image-patch tokens
  │                                # output: length array per sample
  ├─ s3_bin_packing.py             # BFD (Best-Fit-Decreasing) into bins of capacity L
  │                                # output: bin assignment
  └─ s4_bins_to_webdataset.py      # write tar shards + idx + .nv-meta/{dataset.yaml,split.yaml,sample_loader.py}

auto_pipe.sh runs all four stages sequentially on one node for one input file. To use N nodes, split the JSONL into N parts and run auto_pipe.sh independently on each — s3 BFD does NOT shard across nodes, so each node packs its own slice.

auto_pipe.sh 在一台机器上对一个输入文件顺序跑完四个 stage。要用 N 台机器就把 JSONL 切成 N 份,每台独立跑一次 auto_pipe.sh——s3 BFD 不能跨节点共享,所以每台只 pack 自己的那一份。

Key Design Decisions / 关键设计
  • Per-node shard prefix: pass distinct --shard-prefix <name>_<a|b|...> so tar files don't collide on shared NFS
  • Sample class: choose based on data shape — PackedCaptioningSample for image+text packed turns, MultiMixQASample for QA-style. The class controls the auto-generated sample_loader.py
  • --no-npy flag: only set when JSONL has no precomputed patch_positions field. Without --no-npy s2 expects per-sample .npy files; missing files cause noisy warnings AND fall back to slow real-image tokenization
  • Sequence length L: scan token lengths first to confirm drop rate is acceptable
  • Output layout per node: <output_root>/node_<x>/webdataset/{*.tar, *.idx, .nv-meta/}
  • Top-level Metadataset yaml combines all node outputs into one logical dataset

Step-by-step Workflow / 操作步骤

0. Pre-flight: choose L / 预检:选 L

Scan token lengths once on the full JSONL using s2_compute_token_lengths.py (or a quick standalone script). Pick L so that drop rate is acceptable (typically <0.1%).

bash
# Inside container, on any single node
python offline_packing/s2_compute_token_lengths.py \
    --jsonl <path/to/full.jsonl> \
    --tokenizer <path/to/tokenizer> \
    --image-processor Qwen2_5_VLProcessor \
    --factor 48 --min-pixels 3136 --max-pixels 4000000 \
    --output <path/to/token_lens.txt>

# Then quickly inspect distribution (max, p99, count > L) before committing to L
1. Split JSONL across nodes / 切分 JSONL
bash
TOTAL=$(wc -l < full.jsonl)
HALF=$(( (TOTAL + 1) / 2 ))
split -l $HALF -d --additional-suffix=.jsonl full.jsonl part_
# produces part_00.jsonl, part_01.jsonl

For >2 nodes, adjust -l accordingly.

2. Mount NFS on every node / 挂载 NFS

Make sure the same NFS is mounted on every node at the same path so paths in JSONL and outputs match.

3. Launch container on every node / 每台启动容器

Use the project's standard docker image with the repo bind-mounted. Working directory should be the repo root.

4. Run auto_pipe.sh in parallel on each node / 并行启动

On node A:

bash
cd <repo_root>
bash offline_packing/auto_pipe.sh \
    --jsonl <data_root>/part_00.jsonl \
    --tokenizer <tokenizer_path> \
    --image-processor Qwen2_5_VLProcessor \
    --factor 48 --min-pixels 3136 --max-pixels 4000000 \
    --image-root / \
    --sample-class PackedCaptioningSample \
    --shard-prefix <dataset_name>_a \
    --output-dir <output_root>/node_a \
    --seq-len 4096 \
    --no-npy \
    2>&1 | tee <log_dir>/node_a.log

On node B (in parallel):

bash
bash offline_packing/auto_pipe.sh \
    --jsonl <data_root>/part_01.jsonl \
    ... \
    --shard-prefix <dataset_name>_b \
    --output-dir <output_root>/node_b \
    ... \
    2>&1 | tee <log_dir>/node_b.log

[!IMPORTANT]

  • --shard-prefix must differ between nodes so tar filenames don't collide
  • --output-dir must differ between nodes
  • --image-root / if JSONL paths are absolute; otherwise set it to the image root prefix
  • Add --no-npy if the JSONL has no patch_positions field
5. Verify outputs / 验证产物
bash
# Tar count per node should match s4 log
ls <output_root>/node_a/webdataset/*.tar | wc -l
ls <output_root>/node_b/webdataset/*.tar | wc -l

# Bin count + capacity utilization printed by s3
grep -E "(bins|util|efficiency)" <log_dir>/node_a.log

# Inspect one tar to confirm sample schema
mkdir -p /tmp/tar_inspect && cd /tmp/tar_inspect
tar -xf <output_root>/node_a/webdataset/<prefix>-000000.tar
ls | head
python -c "import json; d=json.load(open(open(__import__('glob').glob('*.json')[0]).name)); print(list(d.keys()))"

Expected JSON top-level keys for PackedCaptioningSample: images, prompts, captions, sample_count, patch_positions, timestamp_decimal.

patch_positions=[[""]] is normal under --no-npy — the auto-generated sample_loader.py handles it via sample.get(..., None).

6. Write top-level Metadataset yaml / 写顶层 Metadataset yaml
yaml
__module__: megatron.energon
__class__: Metadataset
splits:
  train:
    datasets:
      - weight: <num_samples_in_part_00>
        path: <output_root>/node_a/webdataset
        subflavors:
          augmentation: false
      - weight: <num_samples_in_part_01>
        path: <output_root>/node_b/webdataset
        subflavors:
          augmentation: false

Notes / 注意:

  • weight is sample-count proportional (use the original input line count of each part, not the bin count)
  • Use absolute paths — Energon resolves them as-is
  • Add val split only if you actually need one; for SFT-only pipelines omit it
  • subflavors is optional; here we mark augmentation: false since data is already packed

Common Pitfalls / 常见坑

Pitfall 1: forgetting --no-npy

Without it s2 hunts for per-sample .npy files and either spams warnings OR falls back to the slow real-image tokenization path. Always check the source JSONL for a patch_positions field first.

Show full SKILL.md (359 more words)Show less
Pitfall 2: same --shard-prefix on multiple nodes

Tar filenames collide on shared NFS, second node overwrites the first. Use distinct suffixes (_a, _b, _n0, _n1, …).

Pitfall 3: starting fresh while old rm -rf still running on NFS

Removing millions of small files from NFS can take 10+ minutes. Don't wait — use a fresh _v2 output dir, and mv (atomic rename on same FS) when done if you want the original name back.

Pitfall 4: du -sh / rm -rf exceeding bash 120s timeout

Run them as nohup ... & and poll the PID, or run them in tmux.

Pitfall 5: container vs host timezone skew

Container time may differ from host time but wall clock is the same. Don't be confused by log timestamps when comparing across docker exec sessions.

Pitfall 6: choosing the wrong --sample-class

The class determines the auto-generated sample_loader.py and how downstream code unpacks tars. Confirm by reading the dataclass file (e.g. aiak_training_llm/data/multimodal/flavors/packed_captioning.py) and matching its fields to your JSONL shape.

Pitfall 7: weight confusion in Metadataset yaml

Use sample counts (or proportional integers), not bin counts. Energon samples each dataset proportional to weight.

Performance Reference / 性能参考

For ~390k samples per node at L=4096 on a multi-core machine with NFS storage:

  • s1 (split + validate): ~30 min (NFS small-file IO bound; gets worse if concurrent rm is running)
  • s2 (token lengths, with --no-npy): ~5 min
  • s3 (BFD bin packing): <1 min
  • s4 (tar writing): ~5 min
  • Total wall time per node: ~40–45 min

Two nodes in parallel ≈ same wall time as one node, so 2× throughput.

Typical packing efficiency at L=4096 with avg ~10 samples/bin: >99% capacity utilization.

Quick Sanity Checklist / 快速自检清单

Before declaring done:

  • Tar count per node matches s4 log
  • s3 log shows >95% capacity utilization
  • At least one tar inspected; JSON keys match the chosen sample class
  • sample_count in tar JSON > 0 (not all 1; means packing actually worked)
  • Top-level dataset.yaml exists with absolute paths and correct weights
  • (Optional) Energon load smoke test passes
  • offline_packing/auto_pipe.sh — pipeline driver
  • offline_packing/s1_split_json_to_samples.py — JSONL → per-sample records
  • offline_packing/s2_compute_token_lengths.py — token length computation
  • offline_packing/s3_bin_packing.py — BFD bin packing
  • offline_packing/s4_bins_to_webdataset.py — tar + idx + .nv-meta writer
  • aiak_training_llm/data/multimodal/flavors/packed_captioning.py — PackedCaptioningSample dataclass
  • aiak_training_llm/data/multimodal/flavors/multi_mix_qa.py — MultiMixQASample dataclass
  • aiak_megatron/examples/multimodal/sft_dataset.yaml — Metadataset yaml template

© EvolvingLMMs-Lab, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .opencode/skills/distributed-offline-packing of EvolvingLMMs-Lab/LLaVA-OneVision-2.

Open the folder on GitHubat commit 6ef16b1

Compare with similar skills

Distributed Offline Packing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Distributed Offline Packing compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Distributed Offline Packing this skillEvolvingLMMs-Lab/LLaVA-OneVision-21.2k—~2.6kAutomated safety check: PassApache-2.0
Book Video Factoryjaxxchen003/book-video-factory104—~2.4kAutomated safety check: PassMIT
Prompts CollectionNorman-bury/research-writing-skill3.4k—~1kAutomated safety check: PassMIT
Translation Diff ExportDevolutions/UniGetUI26k—~1.1kAutomated safety check: PassMIT
Sync Translationssymfony/symfony31k—~1.9kAutomated safety check: PassMIT
Translation Diff ImportDevolutions/UniGetUI26k—~750Automated safety check: PassMIT

Similar skills

  • Book Video Factory

    jaxxchen003/book-video-factory

    Create or operate a portable, auditable Chinese book-review short-video workflow from a clean local workspace.

    104 GitHub stars~2.4k tokensUpdated 1 mo ago
    Writing & ContentAuto-check passed
  • Prompts Collection

    Norman-bury/research-writing-skill

    A skill your agent uses for translation, polishing, or de-AI-ification of academic text - provides ready-to-use prompt templates

    3.4k GitHub stars~1k tokensUpdated 4 mo ago
    Writing & ContentAuto-check passed
  • Translation Diff Export

    Devolutions/UniGetUI

    Compares UniGetUI JSON locale files against English, identifies untranslated or source-changed keys, and generates patch, reference, and handoff files for a target language.

    26k GitHub stars~1.1k tokensUpdated today
    Writing & ContentAuto-check passed
  • Sync Translations

    symfony/symfony

    Synchronize translation catalogs across maintained Symfony branches: find messages that newer branches added to the English catalogs but that are still missing from the oldest maintained branch…

    31k GitHub stars~1.9k tokensUpdated today
    Writing & ContentAuto-check passed
  • Translation Diff Import

    Devolutions/UniGetUI

    Merges translated key-value pairs from a UniGetUI JSON localization patch back into the full language file and validates the merged result.

    26k GitHub stars~750 tokensUpdated today
    Writing & ContentAuto-check passed
  • Translation Diff Translate

    Devolutions/UniGetUI

    Translates a sparse UniGetUI JSON language patch, writes completed entries into the working copy, preserves placeholders and terminology, and prepares the patch for merge-back.

    26k GitHub stars~934 tokensUpdated today
    Writing & ContentAuto-check passed

More from EvolvingLMMs-Lab/LLaVA-OneVision-2

All 8 skills in this repo
  • Commit Message

    EvolvingLMMs-Lab/LLaVA-OneVision-2

    Guide for writing clear, consistent git commit messages following this repository's conventions

    1.2k GitHub stars~1.9k tokensUpdated yesterday
    Auto-check passed
  • Cu Lengths Attention Flow

    EvolvingLMMs-Lab/LLaVA-OneVision-2

    Bilingual guide for understanding how culengths controls attention behavior across ViT and LLM stages, and how patchpositions scope differs between the two

    1.2k GitHub stars~3.1k tokensUpdated yesterday
    Auto-check passed
  • Length Pool Sort Dataset

    EvolvingLMMs-Lab/LLaVA-OneVision-2

    Bilingual guide for understanding LengthPoolSortDataset cross-rank length synchronization mechanism in multi-GPU training

    1.2k GitHub stars~2k tokensUpdated yesterday
    Auto-check passed
  • Llava Onevision2 Consistency

    EvolvingLMMs-Lab/LLaVA-OneVision-2

    Bilingual guide for running and interpreting LLaVA-OneVision2 HF vs Megatron consistency checks across TP and PP settings

    1.2k GitHub stars~4.1k tokensUpdated yesterday
    Auto-check passed
  • Offline Packing Env Vars

    EvolvingLMMs-Lab/LLaVA-OneVision-2

    Bilingual guide for the OFFLINEPACKINGBMR and OFFLINEPACKEDDATA environment variables that control LLaVA-OneVision2 training-side packing — what each gate does, why both must be enabled together…

    1.2k GitHub stars~3.9k tokensUpdated yesterday
    Auto-check passed
  • Merge Ov2

    EvolvingLMMs-Lab/LLaVA-OneVision-2

    Bilingual guide for merging ViT + LLM into LlavaOnevision2 HF checkpoint and validating weight/inference consistency

    1.2k GitHub stars~7.6k tokensUpdated yesterday
    Auto-check passed

Questions about Distributed Offline Packing

What does Distributed Offline Packing do?

Bilingual guide for running offlinepacking/autopipe.sh across multiple nodes to produce padding-free packed WebDataset shards for SFT, with Energon Metadataset assembly. Distributed Offline Packing is an agent skill from EvolvingLMMs-Lab/LLaVA-OneVision-2.

When should I use Distributed Offline Packing?

Distributed Offline Packing fits situations like: tasks that involve Translation.

How do I install Distributed Offline Packing in Claude Code?

Run `npx skills add EvolvingLMMs-Lab/LLaVA-OneVision-2 --skill distributed-offline-packing -a claude-code`. Or copy the skill folder (.opencode/skills/distributed-offline-packing in EvolvingLMMs-Lab/LLaVA-OneVision-2) into .claude/skills/distributed-offline-packing in your project. Claude Code loads it when a task matches its description.

How do I install Distributed Offline Packing in Codex?

Run `npx skills add EvolvingLMMs-Lab/LLaVA-OneVision-2 --skill distributed-offline-packing -a codex`. Or copy the skill folder (.opencode/skills/distributed-offline-packing in EvolvingLMMs-Lab/LLaVA-OneVision-2) into .agents/skills/distributed-offline-packing in your project. Codex loads it when a task matches its description.

Can I use Distributed Offline Packing in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add EvolvingLMMs-Lab/LLaVA-OneVision-2 --skill distributed-offline-packing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/distributed-offline-packing, .gemini/skills/distributed-offline-packing, .github/skills/distributed-offline-packing and .opencode/skills/distributed-offline-packing in your project.

What does Distributed Offline Packing need to run?

Going by SKILL.md and its folder, Distributed Offline Packing needs the command-line tools its instructions call (python and bash). Our summary lists: Python 3; Docker. Compatibility (from SKILL.md): opencode.

Does Distributed Offline Packing access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Distributed Offline Packing safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Distributed Offline Packing use?

Distributed Offline Packing is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Distributed Offline Packing use?

About 2.6k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Distributed Offline Packing?

Skills that share tags, products or a category with Distributed Offline Packing: Book Video Factory (jaxxchen003/book-video-factory, 104 stars), Prompts Collection (Norman-bury/research-writing-skill, 3.4k stars), Translation Diff Export (Devolutions/UniGetUI, 26k stars) and Sync Translations (symfony/symfony, 31k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Distributed Offline Packing?

EvolvingLMMs-Lab (a GitHub organization) maintains it in EvolvingLMMs-Lab/LLaVA-OneVision-2, which has 1,216 GitHub stars. The repository holds 8 skills in this directory. The repository was last updated on October 7, 2026.

Source: EvolvingLMMs-Lab/LLaVA-OneVision-2 on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.