Agent skill

Length Pool Sort Dataset

by EvolvingLMMs-Lab in EvolvingLMMs-Lab/LLaVA-OneVision-2

Bilingual guide for understanding LengthPoolSortDataset cross-rank length synchronization mechanism in multi-GPU training

Apache-2.0Auto-check passedWriting & Content

Install Length Pool Sort Dataset

skills CLI
$ npx skills add EvolvingLMMs-Lab/LLaVA-OneVision-2 --skill length-pool-sort-dataset -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install EvolvingLMMs-Lab/LLaVA-OneVision-2 length-pool-sort-dataset --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/EvolvingLMMs-Lab/LLaVA-OneVision-2.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.opencode/skills/length-pool-sort-dataset .claude/skills/length-pool-sort-dataset && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
length-pool-sort-dataset
GitHub stars
1.2k
Token cost
~2k tokens
SKILL.md length
778 words
Files
1
Skills in repo
8
Repo updated
First seen
Licence
Apache-2.0

At a glance

Bilingual guide for understanding LengthPoolSortDataset cross-rank length synchronization mechanism in multi-GPU training

  • Works in 2 steps: Sort aligns the i-th position across… → Same seed preserves alignment after…
  • Tasks that involve Translation
  • SKILL.md covers Purpose / 用途, Core mechanism / 核心机制, Why it accelerates training /… and pool_size tuning / pool_size 调优, plus 5 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Length Pool Sort Dataset is an agent skill from EvolvingLMMs-Lab/LLaVA-OneVision-2. Bilingual guide for understanding LengthPoolSortDataset cross-rank length synchronization mechanism in multi-GPU training

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts. Compatibility notes: opencode

It sits in Writing & Content, covering Translation. The repository describes itself as: Fully Open Framework for Democratized Multimodal Training. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Translation

Example prompts

  • “/length-pool-sort-dataset”

Requirements

  • Python 3
  • Compatibility (from SKILL.md): opencode

Workflow steps

2 steps, taken from the step headings in SKILL.md.

  1. Sort aligns the i-th position across ranks / 排序对齐各 rank 第 i 个位置
  2. Same seed preserves alignment after shuffle / 同 seed 保持 shuffle 后的对齐

What it can do on your machine

Read from SKILL.md and the folder at commit 6ef16b1. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    opencode

    From compatibility in the SKILL.md frontmatter.

Context cost

Length Pool Sort Dataset loads about 2k tokens when it runs. Until then it costs about 37 tokens; SKILL.md has 778 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~37
When it runs · the whole SKILL.md, loaded when a task matches
~2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from EvolvingLMMs-Lab/LLaVA-OneVision-2 at commit 6ef16b1, republished under its Apache-2.0 licence (© EvolvingLMMs-Lab). 778 words, ~2,008 tokens.

Download SKILL.mdSave it as .claude/skills/length-pool-sort-dataset/SKILL.md (or your agent's skills folder).
name
length-pool-sort-dataset
description
Bilingual guide for understanding LengthPoolSortDataset cross-rank length synchronization mechanism in multi-GPU training
compatibility
opencode
metadata.domain
distributed-training
metadata.framework
megatron-energon
metadata.repo
llava-onevision2

Purpose / 用途

Use this skill when analyzing, debugging, or tuning LengthPoolSortDataset — the cross-rank length synchronization mechanism used in this repository's training pipeline.

在分析、调试或调优本仓库训练 pipeline 中的跨 rank 长度同步机制 LengthPoolSortDataset 时,使用这个 skill。

This skill is specifically for:

  • aiak_training_llm/data/multimodal/length_sort_dataset.py
  • Understanding why training speed improves when length_sort_pool_size > 0
  • Tuning pool_size for optimal multi-GPU efficiency
  • Diagnosing rank synchronization bottlenecks

这个 skill 专门用于:

  • aiak_training_llm/data/multimodal/length_sort_dataset.py
  • 理解为什么 length_sort_pool_size > 0 时训练速度提升
  • 调优 pool_size 以获得最佳多卡效率
  • 排查 rank 间同步瓶颈

Core mechanism / 核心机制

Three-step pipeline / 三步流水线
上游 dataset → 累积 pool_size 个 sample → 按序列长度排序 → 用确定性 seed shuffle → 逐个 yield
python
for batch_idx, sample in enumerate(self.dataset):
    pool.append(sample)
    if len(pool) >= self.pool_size:
        pool.sort(key=self.key_fn)                     # 1. 按长度排序
        shuffle_seed = 42 + batch_idx                   # 2. 确定性 seed
        random.Random(shuffle_seed).shuffle(pool)       # 3. 同 seed shuffle
        for s in pool:
            yield s
        pool.clear()
Pipeline position / 在 pipeline 中的位置
CrudeWebdataset → ShuffleBuffer → cook_crude_sample → encode_sample
    → LengthPoolSortDataset → BatchDataset → EpochizeDataset → LogSampleDataset

Inserted after encode_sample (where total_len / tokens are available) and before BatchDataset.

插在 encode_sample 之后(此时已有 total_len / tokens)、BatchDataset 之前。

Activated by: --length-sort-pool-size N (where N > 0).

通过 --length-sort-pool-size N(N > 0)激活。

Why it accelerates training / 为什么能加速训练

The problem / 问题

In multi-GPU data-parallel training, all ranks must synchronize at each step (gradient all-reduce). If different ranks process samples of very different lengths, fast ranks idle waiting for slow ranks.

多卡数据并行训练中,所有 rank 每步都要同步(梯度 all-reduce)。如果不同 rank 处理的 sample 长度差异很大,快的 rank 空等慢的 rank。

无 pool sort:
  Rank 0 step 100: 长度 200 → 0.5s    ┐
  Rank 1 step 100: 长度 5000 → 3.0s   ├→ 所有 rank 等 3.0s
  Rank 2 step 100: 长度 300 → 0.6s    ┘
The solution / 解决方案

Sort + same-seed shuffle ensures all ranks yield samples of approximately the same length at the same time.

排序 + 同 seed shuffle 保证所有 rank 在同一时刻输出近似相同长度的 sample。

有 pool sort:
  Rank 0 step 100: 长度 ~200 → 0.5s   ┐
  Rank 1 step 100: 长度 ~200 → 0.5s   ├→ 所有 rank 等 0.5s
  Rank 2 step 100: 长度 ~200 → 0.5s   ┘
Why it works — step by step / 为什么有效——逐步分析
Step 1: Sort aligns the i-th position across ranks / 排序对齐各 rank 第 i 个位置

Each rank's data comes from different shards of the same dataset, so length distributions are approximately identical. After sorting, the i-th position in each rank's pool corresponds to the i-th quantile of its length distribution. Similar distributions → similar quantiles → similar lengths.

各 rank 的数据来自同一数据集的不同 shard,长度分布近似相同。排序后,各 rank pool 中第 i 个位置对应各自长度分布的第 i 个分位数。分布相似 → 分位数相似 → 长度相似。

Rank 0 排序后: [100, 102, 105, 108, ..., 4998, 5000]
Rank 1 排序后: [101, 103, 106, 109, ..., 4999, 5001]
Rank 2 排序后: [ 99, 104, 107, 110, ..., 4997, 5002]
Step 2: Same seed preserves alignment after shuffle / 同 seed 保持 shuffle 后的对齐

shuffle_seed = 42 + batch_idx is identical across all ranks → same permutation applied to all pools. Since the i-th position had similar lengths, the shuffled output at the same position still has similar lengths.

shuffle_seed = 42 + batch_idx 对所有 rank 相同 → 相同排列应用于所有 pool。由于第 i 个位置的长度相似,shuffle 后同一位置的输出长度仍然相似。

permutation = [3, 0, 4, 1, 2]:
  Rank 0 输出: [108, 100, 5000, 102, 105]
  Rank 1 输出: [109, 101, 5001, 103, 106]
  Rank 2 输出: [110,  99, 5002, 104, 107]
  → 同一位置长度近似一致
Why both sort AND shuffle are needed / 为什么排序和 shuffle 缺一不可
方案效果
只 sort 不 shuffle所有 rank 先跑短 sample 后跑长 sample。前期 step 极快,后期 step 极慢。训练动态不稳定
sort + shuffle各 step 长度随机但 rank 间一致。step 耗时平稳,rank 间同步开销小
不 sort 不 shuffle各 rank 同一 step 长度随机且不一致,快的等慢的

Sort solves cross-rank synchronization. Shuffle solves temporal uniformity. Both are required.

排序解决跨 rank 同步问题,shuffle 解决时序均匀性问题。二者缺一不可。

pool_size tuning / pool_size 调优

pool_size跨 rank 长度同步精度内存开销首个 pool 输出延迟
小(~100)较差,各 rank 分位数估计不准低低
中(~1000-10000)好,推荐范围中中
大(~全量)完美同步高高
极限 = dataset 大小等价全局排序不实际不实际

Larger pool_size → more accurate quantile estimation across ranks → better synchronization → less idle time.

pool_size 越大 → 各 rank 分位数估计越准 → 同步越好 → 空等越少。

Rule of thumb: pool_size should be significantly larger than batch_size, ideally 10x-100x.

经验法则:pool_size 应远大于 batch_size,理想情况下 10x-100x。

Show full SKILL.md (337 more words)Show less

Multi-worker behavior (num_workers > 1) / 多 worker 行为

When num_workers > 1, each worker runs an independent LengthPoolSortDataset instance with its own pool.

当 num_workers > 1 时,每个 worker 运行独立的 LengthPoolSortDataset 实例,各自维护独立 pool。

  • Correctness: No issue. Each worker sorts its own shard subset independently.

  • Synchronization effect: Diluted. DataLoader interleaves outputs from multiple workers, partially disrupting the within-pool length ordering.

  • Recommendation: num_workers=1 gives the strongest synchronization effect. If num_workers > 1 is needed for I/O throughput, increase pool_size proportionally.

  • 正确性:无问题。每个 worker 独立排序自己的 shard 子集。

  • 同步效果:被稀释。DataLoader 交替取多个 worker 的输出,部分打乱 pool 内的长度顺序。

  • 建议:num_workers=1 同步效果最强。如果需要 num_workers > 1 提升 I/O 吞吐,可按比例增大 pool_size。

Checkpoint resume caveat / checkpoint 恢复注意事项

save_state / restore_state delegate directly to the upstream dataset. The pool's internal state (accumulated but not yet yielded samples) is NOT saved.

save_state / restore_state 直接委托给上游 dataset。pool 内部状态(已累积但未 yield 的 sample)不会被保存。

On resume, up to pool_size - 1 samples may be re-ordered differently or skipped. This is generally negligible for large datasets.

恢复时,最多 pool_size - 1 个 sample 可能会被不同地排序或跳过。对于大数据集来说通常可以忽略。

key_fn / 排序键

python
key_fn=lambda s: getattr(s, "total_len", len(getattr(s, "tokens")))

Prefers total_len attribute (set by encode_sample), falls back to len(tokens).

优先使用 total_len 属性(由 encode_sample 设置),回退到 len(tokens)。

What to check during debugging / 调试时要检查什么

  • Is --length-sort-pool-size actually set and > 0?

  • What is the actual length distribution of the dataset? (High variance → more benefit from pool sort)

  • Are all ranks receiving data from the same underlying dataset with similar length distributions?

  • Is num_workers > 1 diluting the synchronization effect?

  • After enabling, did step time variance across steps decrease?

  • Is memory usage acceptable for the chosen pool_size?

  • --length-sort-pool-size 是否真正设置且 > 0?

  • 数据集的实际长度分布如何?(方差越大 → pool sort 收益越大)

  • 所有 rank 是否从同一底层数据集接收长度分布相似的数据?

  • num_workers > 1 是否稀释了同步效果?

  • 启用后,各 step 的 step time 方差是否降低了?

  • 所选 pool_size 的内存占用是否可接受?

Expected outputs when using this skill / 使用本 skill 时的期望输出

When asked to analyze a training speed issue related to length sorting, return:

当被要求分析与长度排序相关的训练速度问题时,应返回:

  1. Whether LengthPoolSortDataset is active and with what pool_size

  2. The dataset's length distribution characteristics (high/low variance)

  3. The relationship between pool_size, batch_size, and num_workers

  4. Whether cross-rank synchronization is the bottleneck

  5. Recommended pool_size adjustment if needed

  6. LengthPoolSortDataset 是否激活,pool_size 是多少

  7. 数据集的长度分布特征(高/低方差)

  8. pool_size、batch_size、num_workers 之间的关系

  9. 跨 rank 同步是否是瓶颈

  10. 如需要,推荐的 pool_size 调整

© EvolvingLMMs-Lab, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .opencode/skills/length-pool-sort-dataset of EvolvingLMMs-Lab/LLaVA-OneVision-2.

Open the folder on GitHubat commit 6ef16b1

Compare with similar skills

Length Pool Sort Dataset next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Length Pool Sort Dataset compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Length Pool Sort Dataset this skillEvolvingLMMs-Lab/LLaVA-OneVision-21.2k—~2kAutomated safety check: PassApache-2.0
Translation Diff ExportDevolutions/UniGetUI26k—~1.1kAutomated safety check: PassMIT
Sync Translationssymfony/symfony31k—~1.9kAutomated safety check: PassMIT
Translation Diff ImportDevolutions/UniGetUI26k—~750Automated safety check: PassMIT
Translation Diff TranslateDevolutions/UniGetUI26k—~934Automated safety check: PassMIT
Generate Translationspayloadcms/payload45k—~1.1kAutomated safety check: PassMIT

Similar skills

  • Translation Diff Export

    Devolutions/UniGetUI

    Compares UniGetUI JSON locale files against English, identifies untranslated or source-changed keys, and generates patch, reference, and handoff files for a target language.

    26k GitHub stars~1.1k tokensUpdated yesterday
    Writing & ContentAuto-check passed
  • Sync Translations

    symfony/symfony

    Synchronize translation catalogs across maintained Symfony branches: find messages that newer branches added to the English catalogs but that are still missing from the oldest maintained branch…

    31k GitHub stars~1.9k tokensUpdated today
    Writing & ContentAuto-check passed
  • Translation Diff Import

    Devolutions/UniGetUI

    Merges translated key-value pairs from a UniGetUI JSON localization patch back into the full language file and validates the merged result.

    26k GitHub stars~750 tokensUpdated yesterday
    Writing & ContentAuto-check passed
  • Translation Diff Translate

    Devolutions/UniGetUI

    Translates a sparse UniGetUI JSON language patch, writes completed entries into the working copy, preserves placeholders and terminology, and prepares the patch for merge-back.

    26k GitHub stars~934 tokensUpdated yesterday
    Writing & ContentAuto-check passed
  • Generate Translations

    payloadcms/payload

    A skill your agent uses when new translation keys are added to packages to generate new translations strings

    45k GitHub stars~1.1k tokensUpdated today
    Writing & ContentAuto-check passed
  • Drives long-form fiction, scripts, storyboards, interactive films and long-document translation through InkOS, with every change made by a typed action.

    10k GitHub starsUsed in 1 repo~1.1k tokens
    Writing & ContentAuto-check passed

More from EvolvingLMMs-Lab/LLaVA-OneVision-2

All 8 skills in this repo
  • Commit Message

    EvolvingLMMs-Lab/LLaVA-OneVision-2

    Guide for writing clear, consistent git commit messages following this repository's conventions

    1.2k GitHub stars~1.9k tokensUpdated today
    Auto-check passed
  • Cu Lengths Attention Flow

    EvolvingLMMs-Lab/LLaVA-OneVision-2

    Bilingual guide for understanding how culengths controls attention behavior across ViT and LLM stages, and how patchpositions scope differs between the two

    1.2k GitHub stars~3.1k tokensUpdated today
    Auto-check passed
  • Distributed Offline Packing

    EvolvingLMMs-Lab/LLaVA-OneVision-2

    Bilingual guide for running offlinepacking/autopipe.sh across multiple nodes to produce padding-free packed WebDataset shards for SFT, with Energon Metadataset assembly

    1.2k GitHub stars~2.6k tokensUpdated today
    Auto-check passed
  • Llava Onevision2 Consistency

    EvolvingLMMs-Lab/LLaVA-OneVision-2

    Bilingual guide for running and interpreting LLaVA-OneVision2 HF vs Megatron consistency checks across TP and PP settings

    1.2k GitHub stars~4.1k tokensUpdated today
    Auto-check passed
  • Offline Packing Env Vars

    EvolvingLMMs-Lab/LLaVA-OneVision-2

    Bilingual guide for the OFFLINEPACKINGBMR and OFFLINEPACKEDDATA environment variables that control LLaVA-OneVision2 training-side packing — what each gate does, why both must be enabled together…

    1.2k GitHub stars~3.9k tokensUpdated today
    Auto-check passed
  • Merge Ov2

    EvolvingLMMs-Lab/LLaVA-OneVision-2

    Bilingual guide for merging ViT + LLM into LlavaOnevision2 HF checkpoint and validating weight/inference consistency

    1.2k GitHub stars~7.6k tokensUpdated today
    Auto-check passed

Questions about Length Pool Sort Dataset

What does Length Pool Sort Dataset do?

Bilingual guide for understanding LengthPoolSortDataset cross-rank length synchronization mechanism in multi-GPU training. Length Pool Sort Dataset is an agent skill from EvolvingLMMs-Lab/LLaVA-OneVision-2.

When should I use Length Pool Sort Dataset?

Length Pool Sort Dataset fits situations like: tasks that involve Translation.

How do I install Length Pool Sort Dataset in Claude Code?

Run `npx skills add EvolvingLMMs-Lab/LLaVA-OneVision-2 --skill length-pool-sort-dataset -a claude-code`. Or copy the skill folder (.opencode/skills/length-pool-sort-dataset in EvolvingLMMs-Lab/LLaVA-OneVision-2) into .claude/skills/length-pool-sort-dataset in your project. Claude Code loads it when a task matches its description.

How do I install Length Pool Sort Dataset in Codex?

Run `npx skills add EvolvingLMMs-Lab/LLaVA-OneVision-2 --skill length-pool-sort-dataset -a codex`. Or copy the skill folder (.opencode/skills/length-pool-sort-dataset in EvolvingLMMs-Lab/LLaVA-OneVision-2) into .agents/skills/length-pool-sort-dataset in your project. Codex loads it when a task matches its description.

Can I use Length Pool Sort Dataset in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add EvolvingLMMs-Lab/LLaVA-OneVision-2 --skill length-pool-sort-dataset -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/length-pool-sort-dataset, .gemini/skills/length-pool-sort-dataset, .github/skills/length-pool-sort-dataset and .opencode/skills/length-pool-sort-dataset in your project.

What does Length Pool Sort Dataset need to run?

SKILL.md names no scripts, command-line tools or credentials: Length Pool Sort Dataset is instructions for the agent only. Our summary lists: Python 3. Compatibility (from SKILL.md): opencode.

Does Length Pool Sort Dataset access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Length Pool Sort Dataset safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Length Pool Sort Dataset use?

Length Pool Sort Dataset is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Length Pool Sort Dataset use?

About 2k tokens (SKILL.md is roughly 8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Length Pool Sort Dataset?

Skills that share tags, products or a category with Length Pool Sort Dataset: Translation Diff Export (Devolutions/UniGetUI, 26k stars), Sync Translations (symfony/symfony, 31k stars), Translation Diff Import (Devolutions/UniGetUI, 26k stars) and Translation Diff Translate (Devolutions/UniGetUI, 26k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Length Pool Sort Dataset?

EvolvingLMMs-Lab (a GitHub organization) maintains it in EvolvingLMMs-Lab/LLaVA-OneVision-2, which has 1,216 GitHub stars. The repository holds 8 skills in this directory. The repository was last updated on October 7, 2026.

Source: EvolvingLMMs-Lab/LLaVA-OneVision-2 on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.