Agent skill

Offline Packing Env Vars

by EvolvingLMMs-Lab in EvolvingLMMs-Lab/LLaVA-OneVision-2

Bilingual guide for the OFFLINEPACKINGBMR and OFFLINEPACKEDDATA environment variables that control LLaVA-OneVision2 training-side packing — what each gate does, why both must be enabled together…

Apache-2.0Auto-check passedWriting & Content

Install Offline Packing Env Vars

skills CLI
$ npx skills add EvolvingLMMs-Lab/LLaVA-OneVision-2 --skill offline-packing-env-vars -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install EvolvingLMMs-Lab/LLaVA-OneVision-2 offline-packing-env-vars --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/EvolvingLMMs-Lab/LLaVA-OneVision-2.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.opencode/skills/offline-packing-env-vars .claude/skills/offline-packing-env-vars && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
offline-packing-env-vars
GitHub stars
1.2k
Token cost
~3.9k tokens
SKILL.md length
1,231 words
Files
1
Skills in repo
8
Repo updated
First seen
Licence
Apache-2.0

At a glance

Bilingual guide for the OFFLINEPACKINGBMR and OFFLINEPACKEDDATA environment variables that control LLaVA-OneVision2 training-side packing — what each gate does, why both must be enabled together…

  • Works in 6 steps: grep -n… → Add a one-shot print in… → Add a print in… → …
  • Tasks that involve Translation
  • SKILL.md covers Purpose / 用途, TL;DR / 一句话总结, The Three Env Vars / 三个环境变量真相表 and Two-Stage Gate Architecture /…, plus 7 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Offline Packing Env Vars is an agent skill from EvolvingLMMs-Lab/LLaVA-OneVision-2. Bilingual guide for the OFFLINEPACKINGBMR and OFFLINEPACKEDDATA environment variables that control LLaVA-OneVision2 training-side packing — what each gate does, why both must be enabled together, MBS=1 requirement, and the dead OFFLINEPACKINGVQA branch

Its SKILL.md is about 3.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts. Compatibility notes: opencode

It sits in Writing & Content, covering Translation and Secrets management. The repository describes itself as: Fully Open Framework for Democratized Multimodal Training. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Translation
  • Tasks that involve Secrets management

Example prompts

  • “/offline-packing-env-vars”

Requirements

  • Python 3
  • Compatibility (from SKILL.md): opencode

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. grep -n 'OFFLINE_PACKING_BMR|OFFLINE_PACKED_DATA' your_script.sh — both should be '1'.
  2. Add a one-shot print in task_encoder.batch() after line 365: print('cu_lengths.shape:', cu_lengths.shape). Expect [1, P+1] with P >= 2. If…
  3. Add a print in pretrain_llava_onevision2.py after line 168: print('packed_seq_params:', packed_seq_params). Should be a real…
  4. Confirm MBS=1 in the shell (--micro-batch-size 1). Otherwise the assert at line 157 fires and you wouldn't be reading this.
  5. Confirm dataset is actually packed: cat $DATA_PATH/.../webdataset/.nv-meta/.info.yaml — look for shard structure produced by auto_pipe.sh…
  6. Do not add OFFLINE_PACKING_VQA=1 thinking it helps. It does nothing in this codebase.

What it can do on your machine

Read from SKILL.md and the folder at commit 6ef16b1. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python and bash).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    opencode

    From compatibility in the SKILL.md frontmatter.

Context cost

Offline Packing Env Vars loads about 3.9k tokens when it runs. Until then it costs about 71 tokens; SKILL.md has 1,231 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~71
When it runs · the whole SKILL.md, loaded when a task matches
~3.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from EvolvingLMMs-Lab/LLaVA-OneVision-2 at commit 6ef16b1, republished under its Apache-2.0 licence (© EvolvingLMMs-Lab). 1,231 words, ~3,874 tokens.

Download SKILL.mdSave it as .claude/skills/offline-packing-env-vars/SKILL.md (or your agent's skills folder).
name
offline-packing-env-vars
description
Bilingual guide for the OFFLINE_PACKING_BMR and OFFLINE_PACKED_DATA environment variables that control LLaVA-OneVision2 training-side packing — what each gate does, why both must be enabled together, MBS=1 requirement, and the dead OFFLINE_PACKING_VQA branch
compatibility
opencode
metadata.domain
training-pipeline
metadata.framework
llava-onevision2
metadata.repo
llava-onevision2

Purpose / 用途

Use this skill when you set up or debug training-side sample packing for LLaVA-OneVision2 — i.e. when you need to decide which env vars to export in a training shell script (Stage-1 / Stage-1.5 / Stage-2) and want to understand why both OFFLINE_PACKING_BMR and OFFLINE_PACKED_DATA must be 1 to actually get padding-free attention.

在配置或调试 LLaVA-OneVision2 训练侧的样本 packing 时使用——比如要决定在训练 shell 脚本(Stage-1 / Stage-1.5 / Stage-2)中导出哪些环境变量,以及为什么必须 OFFLINE_PACKING_BMR=1 和 OFFLINE_PACKED_DATA=1 同时打开才能真正获得 padding-free 的 attention。

This skill is specifically for:

  • Choosing the correct env var combination in training scripts
  • Diagnosing cross-sample attention leakage in packed runs
  • Understanding why cu_lengths is a dummy [[0]] in some runs and a real [B, P+1] tensor in others
  • Avoiding the well-known OFFLINE_PACKING_VQA red herring (it is dead code)

Companion skill: cu-lengths-attention-flow covers the consumer side (how cu_lengths is fed into ViT/LLM attention). This skill covers the producer + gate side.

姊妹 skill:cu-lengths-attention-flow 讲消费端(cu_lengths 如何送入 ViT/LLM attention)。本 skill 讲生产端 + 开关。


TL;DR / 一句话总结

For packed training to work end-to-end, both env vars must be 1:

bash
export OFFLINE_PACKING_BMR='1'   # data-layer gate: build real cu_lengths
export OFFLINE_PACKED_DATA='1'   # batch-layer gate: forward real cu_lengths to model

Setting only one is a silent bug. OFFLINE_PACKING_VQA is dead code; do not rely on it.


The Three Env Vars / 三个环境变量真相表

Env varStatusDefaultRead atEffect
OFFLINE_PACKING_BMRALIVE0aiak_training_llm/data/multimodal/task_encoder.py:194Inside PackedCaptioningSample handling, unroll each packed entry into a MultiMixQASample (BMR-style, with full prompt/caption messages). When 0, falls through to the legacy CaptioningSample branch which loses the multi-turn structure.
OFFLINE_PACKED_DATAALIVE0aiak_training_llm/data/multimodal/task_encoder.py:363Inside batch(), replace dummy cu_lengths = [[0]] with the real per-sample s.cu_lengths stacked across the batch. Without this, the consumer side cannot construct PackedSeqParams.
OFFLINE_PACKING_VQADEADn/anowhere in aiak_training_llm/Mentioned in README + several legacy shells under examples/llava_onevision1_5/ and examples/llava_onevision2/quick_start_video_2b/, but no source file reads it. Setting it has zero runtime effect. Treat as documentation noise.

💡 The OFFLINE_PACKING_VQA red herring is the #1 source of confusion. Newcomers see it in shell scripts and assume it controls VQA packing. It does not. There is no third packing branch in task_encoder.py — only the BMR branch and the legacy captioning fallback.

💡 OFFLINE_PACKING_VQA 这个红鲱鱼是头号困惑源。新人在 shell 脚本里看到它,以为它控制 VQA packing。并不。task_encoder.py 里没有第三个 packing 分支——只有 BMR 分支和老的 captioning fallback。


Two-Stage Gate Architecture / 两段式 Gate 架构

Packing in this codebase is split into two orthogonal gates that must both fire. Understanding this is the whole point of the skill.

本仓库的 packing 拆成两个正交的 gate,必须都触发。理解这一点就是本 skill 的核心。

Gate 1 — Data Layer (OFFLINE_PACKING_BMR)

Where: aiak_training_llm/data/multimodal/task_encoder.py, inside the PackedCaptioningSample branch of the encoder dispatch (encode_sample ~line 186).

What it does:

  • For each entry inside the packed sample (for idx in range(n_orig_sample):), if OFFLINE_PACKING_BMR == 1, it builds a MultiMixQASample carrying the full chat-format messages ({role: user, content: prompt}, {role: assistant, content: caption}) and routes it through encode_multi_mix_qa().
  • If OFFLINE_PACKING_BMR != 1, it falls back to a plain CaptioningSample and encode_captioning() — losing the multi-turn / multi-image structure required for SFT.
  • After the per-entry loop, regardless of the BMR flag, it calls self.pack_selected_samples(l_Qwen2VLImageTaskSample) (line 277), which constructs the per-sub-sample cumulative lengths cu_lengths = [0, len₁, len₁+len₂, ...] and attaches them to the resulting ImageTaskSamplePacked (line 473).

Net effect: enables the correct per-sub-sample encoding and produces real s.cu_lengths on each sample.

作用:启用正确的逐子样本编码,并在每个样本上产出真正的 s.cu_lengths。

⚠️ Even with BMR off, pack_selected_samples still attaches a cu_lengths tensor to the sample. But the sub-samples were encoded via the wrong path (legacy captioning), so the resulting boundaries don't match what the LLM actually sees. BMR off + PACKED_DATA on is a hidden corruption, not just a missing-feature.

⚠️ 即使 BMR 关掉,pack_selected_samples 仍然会给样本挂上 cu_lengths 张量。但子样本走的是错误的编码路径(老 captioning),结果 boundary 和 LLM 实际看到的 token 序列对不上。BMR 关 + PACKED_DATA 开是隐性数据损坏,不只是缺特性。

Gate 2 — Batch Layer (OFFLINE_PACKED_DATA)

Where: aiak_training_llm/data/multimodal/task_encoder.py:359-365, inside batch() (the collate function).

What it does:

python
# Cumulative sample lengths are needed for packing, otherwise use dummy values.
cu_lengths = torch.tensor([[0]], dtype=torch.int32)
max_lengths = torch.tensor([[0]], dtype=torch.int32)

if self.is_packing_enabled or int(os.environ.get("OFFLINE_PACKED_DATA", 0)) == 1:
    cu_lengths = torch.stack([s.cu_lengths for s in samples])
    max_lengths = torch.tensor([s.max_length for s in samples], dtype=torch.int32)
  • Default: emit a dummy cu_lengths of shape [1, 1] containing only [[0]].
  • When OFFLINE_PACKED_DATA == 1 (or the energon online-packing flag is set): stack the real per-sample cu_lengths produced by Gate 1 into shape [B, P+1].

Net effect: decides whether the consumer (model forward) sees real packing offsets or a dummy that says "no packing".

作用:决定消费端(模型 forward)看到的是真实的 packing 偏移,还是一个表示"没有 packing"的 dummy。

Why both gates must fire / 为什么必须两个都开

The consumer side at aiak_training_llm/train/pretrain/pretrain_llava_onevision2.py:153-168:

python
packed_seq_params = None
...
if cu_lengths.shape == torch.Size([1, 1]):
    pass                        # treat as not packed
else:
    assert cu_lengths.shape[0] == 1, "micro-batch-size must be 1 for packing"
    packed_seq_params = PackedSeqParams(
        qkv_format="thd",
        cu_seqlens_q=cu_lengths[0],
        cu_seqlens_kv=cu_lengths[0],
        ...
    )

So:

BMRPACKED_DATAResult
00No packing. Each sample treated independently. Slow but correct (if data is unpacked).
10SILENT BUG. Data is encoded as packed sub-samples (BMR), cu_lengths is built, but batch() discards it as dummy [[0]]. Consumer sees shape == [1,1] → packed_seq_params = None → flash-attn applies a single causal mask across the entire packed sequence → cross-sub-sample attention leakage. Loss looks fine; model silently learns wrong attention.
01HIDDEN CORRUPTION. Sub-samples encoded via legacy path, boundaries in cu_lengths don't align with token sequence. Consumer applies varlen attention with wrong offsets.
11CORRECT. BMR encodes properly, PACKED_DATA forwards the real offsets, consumer builds PackedSeqParams, flash-attn applies per-sub-sample causal mask via cu_seqlens_q/kv.

🔥 The "BMR=1, PACKED_DATA=0" footgun is the most dangerous combination. Training does not crash. Loss curves look reasonable. But every sub-sample in a packed sequence can attend to every other sub-sample's prefix. Use this skill's TL;DR snippet to avoid it.

🔥 "BMR=1, PACKED_DATA=0" 这个组合最危险。训练不会挂,loss 曲线看着也正常。但 packed 序列里每个子样本都能 attend 到别的子样本的 prefix。用本 skill 顶部的 TL;DR 片段避开它。


Show full SKILL.md (420 more words)Show less

MBS=1 Hard Requirement / MBS=1 硬性要求

pretrain_llava_onevision2.py:157:

python
assert cu_lengths.shape[0] == 1, "micro-batch-size must be 1 for packing"

When packing is on, cu_lengths has shape [B, P+1] where B = micro_batch_size and P = number of sub-samples in a packed sequence. The current PackedSeqParams construction only handles B=1 (it indexes cu_lengths[0]). Therefore:

  • --micro-batch-size 1 is mandatory for any packed training run.
  • Increase throughput via --global-batch-size (gradient accumulation), pipeline parallelism, or longer --seq-length, not via MBS.
  • If you forget, the assert fires immediately on the first batch.

打开 packing 时,cu_lengths 形状是 [B, P+1],B = micro batch size,P = 一个 packed 序列里的子样本数。当前 PackedSeqParams 构造只处理 B=1(取 cu_lengths[0])。所以:

  • 打包训练必须 --micro-batch-size 1。
  • 想提吞吐就调 --global-batch-size(梯度累积)、PP 并行、或更长的 --seq-length,不要调 MBS。
  • 忘了的话第一个 batch 就 assert 挂掉。

End-to-End Flow / 端到端流程图

┌─────────────────────────────────────────────────────────────────┐
│ Offline preprocessing (auto_pipe.sh, separate skill)            │
│   Produces WebDataset shards with PackedCaptioningSample format │
└──────────────────────────────┬──────────────────────────────────┘
                               │
                    Energon dataloader yields PackedCaptioningSample
                               │
                               ▼
┌─────────────────────────────────────────────────────────────────┐
│ task_encoder.encode_sample()                                    │
│   if OFFLINE_PACKING_BMR == 1:                  ◄── GATE 1     │
│     for each sub-sample → MultiMixQASample → encode_multi_mix_qa │
│   else:                                                         │
│     for each sub-sample → CaptioningSample → encode_captioning  │
│   pack_selected_samples(l_samples)                              │
│     → ImageTaskSamplePacked with cu_lengths=[0,L1,L1+L2,...]   │
└──────────────────────────────┬──────────────────────────────────┘
                               │
                               ▼
┌─────────────────────────────────────────────────────────────────┐
│ task_encoder.batch()                                            │
│   if is_packing_enabled or OFFLINE_PACKED_DATA==1:  ◄── GATE 2 │
│     cu_lengths = stack([s.cu_lengths for s in samples])         │
│   else:                                                         │
│     cu_lengths = [[0]]    # dummy, signals "not packed"         │
└──────────────────────────────┬──────────────────────────────────┘
                               │ batch dict broadcast via tensor_parallel
                               ▼
┌─────────────────────────────────────────────────────────────────┐
│ pretrain_llava_onevision2.get_batch_on_this_tp_rank()           │
│   if cu_lengths.shape == [1,1]: packed_seq_params = None        │
│   else:                                                         │
│     assert cu_lengths.shape[0] == 1   # MBS=1 required          │
│     packed_seq_params = PackedSeqParams(                        │
│       qkv_format="thd",                                         │
│       cu_seqlens_q=cu_lengths[0],                               │
│       cu_seqlens_kv=cu_lengths[0], ...)                         │
└──────────────────────────────┬──────────────────────────────────┘
                               │
                               ▼
                Model forward → flash-attn varlen
                (see cu-lengths-attention-flow skill)

Recipe: Correct Stage-N Script Snippet / 正确的训练脚本片段

bash
# ───────────────────────────────────────────────────────────
# Packing env vars — both REQUIRED for padding-free training
# Set both to '1' when DATA_PATH points to offline-packed shards
#   (PackedCaptioningSample format, e.g. produced by auto_pipe.sh)
# Leave both as '0' (or unset) for unpacked datasets.
# Mixed states are silent bugs — see offline-packing-env-vars skill.
# ───────────────────────────────────────────────────────────
export OFFLINE_PACKING_BMR='1'
export OFFLINE_PACKED_DATA='1'

# Hard requirement when packing is on
MBS=1
# Throughput knobs: GBS via grad-accum, longer SEQ_LEN, more PP — not MBS

For an A/B control run that uses the same packed dataset but disables packing semantics (to measure the leakage cost), set both to '0'. Setting only BMR=1 or only PACKED_DATA=1 is not a valid configuration — it is a bug.

如果想做 A/B 对照,用同一份 packed 数据但关闭 packing 语义(为了量化 leakage 损失),两个都设 '0'。只开一个不是合法配置,是 bug。


Concrete Stage-1 A/B Pair (this repo) / 本仓库的 Stage-1 A/B 对照

  • examples/llava_onevision2/quick_start_4b/stage_1_alignment_p16m3_packed.sh — production: BMR=1, PACKED_DATA=1.
  • examples/llava_onevision2/quick_start_4b/stage_1_alignment_p16m3_packed_bmr_only.sh — A/B control: BMR=1, PACKED_DATA=0. Note: this is the dangerous combo described above; it is named bmr_only deliberately to study the leakage effect, not as a recommended setting.

If you copy _bmr_only.sh for a real production run, you will get cross-sub-sample attention leakage. Always confirm intent.

如果你把 _bmr_only.sh 拷去做正式训练,就会得到跨子样本 attention leakage。务必确认是有意为之。


Diagnostics / 排查清单

If your packed training looks "off" (loss too low / too smooth / model overfits prefixes):

  1. grep -n 'OFFLINE_PACKING_BMR\|OFFLINE_PACKED_DATA' your_script.sh — both should be '1'.
  2. Add a one-shot print in task_encoder.batch() after line 365: print('cu_lengths.shape:', cu_lengths.shape). Expect [1, P+1] with P >= 2. If you see [1, 1], Gate 2 is closed.
  3. Add a print in pretrain_llava_onevision2.py after line 168: print('packed_seq_params:', packed_seq_params). Should be a real PackedSeqParams, not None.
  4. Confirm MBS=1 in the shell (--micro-batch-size 1). Otherwise the assert at line 157 fires and you wouldn't be reading this.
  5. Confirm dataset is actually packed: cat $DATA_PATH/.../webdataset/.nv-meta/.info.yaml — look for shard structure produced by auto_pipe.sh (PackedCaptioningSample).
  6. Do not add OFFLINE_PACKING_VQA=1 thinking it helps. It does nothing in this codebase.

Cross-References / 交叉引用

  • Producer pipeline (how the packed shards are built): distributed-offline-packing skill.
  • Consumer attention semantics (how cu_lengths is interpreted by ViT and LLM): cu-lengths-attention-flow skill.
  • Dataloader length-balancing across ranks: length-pool-sort-dataset skill.

Source File Index / 源文件索引

FileLinesWhat
aiak_training_llm/data/multimodal/task_encoder.py186-279PackedCaptioningSample branch + Gate 1 (OFFLINE_PACKING_BMR)
aiak_training_llm/data/multimodal/task_encoder.py359-365batch() Gate 2 (OFFLINE_PACKED_DATA)
aiak_training_llm/data/multimodal/task_encoder.py401-477pack_selected_samples — builds real cu_lengths
aiak_training_llm/train/pretrain/pretrain_llava_onevision2.py145-168Consumer: cu_lengths.shape check + PackedSeqParams construction + MBS=1 assert
aiak_training_llm/train/pretrain/pretrain_llava_onevision2.py171-207SP padding for packed_seq_params (TP/SP-only path)

© EvolvingLMMs-Lab, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .opencode/skills/offline-packing-env-vars of EvolvingLMMs-Lab/LLaVA-OneVision-2.

Open the folder on GitHubat commit 6ef16b1

Compare with similar skills

Offline Packing Env Vars next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Offline Packing Env Vars compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Offline Packing Env Vars this skillEvolvingLMMs-Lab/LLaVA-OneVision-21.2k—~3.9kAutomated safety check: PassApache-2.0
Staticphp Documentation Synccrazywhalecc/static-php-cli1.9k—~2.2kAutomated safety check: PassMIT
Translation Diff ExportDevolutions/UniGetUI26k—~1.1kAutomated safety check: PassMIT
Sync Translationssymfony/symfony31k—~1.9kAutomated safety check: PassMIT
Translation Diff ImportDevolutions/UniGetUI26k—~750Automated safety check: PassMIT
Translation Diff TranslateDevolutions/UniGetUI26k—~934Automated safety check: PassMIT

Similar skills

  • Staticphp Documentation Sync

    crazywhalecc/static-php-cli

    Synchronize bilingual documentation when StaticPHP v3 user-facing or developer-facing documentation must change.

    1.9k GitHub stars~2.2k tokensUpdated yesterday
    DevelopmentAuto-check passed
  • Translation Diff Export

    Devolutions/UniGetUI

    Compares UniGetUI JSON locale files against English, identifies untranslated or source-changed keys, and generates patch, reference, and handoff files for a target language.

    26k GitHub stars~1.1k tokensUpdated today
    Writing & ContentAuto-check passed
  • Sync Translations

    symfony/symfony

    Synchronize translation catalogs across maintained Symfony branches: find messages that newer branches added to the English catalogs but that are still missing from the oldest maintained branch…

    31k GitHub stars~1.9k tokensUpdated today
    Writing & ContentAuto-check passed
  • Translation Diff Import

    Devolutions/UniGetUI

    Merges translated key-value pairs from a UniGetUI JSON localization patch back into the full language file and validates the merged result.

    26k GitHub stars~750 tokensUpdated today
    Writing & ContentAuto-check passed
  • Translation Diff Translate

    Devolutions/UniGetUI

    Translates a sparse UniGetUI JSON language patch, writes completed entries into the working copy, preserves placeholders and terminology, and prepares the patch for merge-back.

    26k GitHub stars~934 tokensUpdated today
    Writing & ContentAuto-check passed
  • Generate Translations

    payloadcms/payload

    A skill your agent uses when new translation keys are added to packages to generate new translations strings

    45k GitHub stars~1.1k tokensUpdated today
    Writing & ContentAuto-check passed

More from EvolvingLMMs-Lab/LLaVA-OneVision-2

All 8 skills in this repo
  • Commit Message

    EvolvingLMMs-Lab/LLaVA-OneVision-2

    Guide for writing clear, consistent git commit messages following this repository's conventions

    1.2k GitHub stars~1.9k tokensUpdated yesterday
    Auto-check passed
  • Cu Lengths Attention Flow

    EvolvingLMMs-Lab/LLaVA-OneVision-2

    Bilingual guide for understanding how culengths controls attention behavior across ViT and LLM stages, and how patchpositions scope differs between the two

    1.2k GitHub stars~3.1k tokensUpdated yesterday
    Auto-check passed
  • Distributed Offline Packing

    EvolvingLMMs-Lab/LLaVA-OneVision-2

    Bilingual guide for running offlinepacking/autopipe.sh across multiple nodes to produce padding-free packed WebDataset shards for SFT, with Energon Metadataset assembly

    1.2k GitHub stars~2.6k tokensUpdated yesterday
    Auto-check passed
  • Length Pool Sort Dataset

    EvolvingLMMs-Lab/LLaVA-OneVision-2

    Bilingual guide for understanding LengthPoolSortDataset cross-rank length synchronization mechanism in multi-GPU training

    1.2k GitHub stars~2k tokensUpdated yesterday
    Auto-check passed
  • Llava Onevision2 Consistency

    EvolvingLMMs-Lab/LLaVA-OneVision-2

    Bilingual guide for running and interpreting LLaVA-OneVision2 HF vs Megatron consistency checks across TP and PP settings

    1.2k GitHub stars~4.1k tokensUpdated yesterday
    Auto-check passed
  • Merge Ov2

    EvolvingLMMs-Lab/LLaVA-OneVision-2

    Bilingual guide for merging ViT + LLM into LlavaOnevision2 HF checkpoint and validating weight/inference consistency

    1.2k GitHub stars~7.6k tokensUpdated yesterday
    Auto-check passed

Questions about Offline Packing Env Vars

What does Offline Packing Env Vars do?

Bilingual guide for the OFFLINEPACKINGBMR and OFFLINEPACKEDDATA environment variables that control LLaVA-OneVision2 training-side packing — what each gate does, why both must be enabled together…. Offline Packing Env Vars is an agent skill from EvolvingLMMs-Lab/LLaVA-OneVision-2.

When should I use Offline Packing Env Vars?

Offline Packing Env Vars fits situations like: tasks that involve Translation; tasks that involve Secrets management.

How do I install Offline Packing Env Vars in Claude Code?

Run `npx skills add EvolvingLMMs-Lab/LLaVA-OneVision-2 --skill offline-packing-env-vars -a claude-code`. Or copy the skill folder (.opencode/skills/offline-packing-env-vars in EvolvingLMMs-Lab/LLaVA-OneVision-2) into .claude/skills/offline-packing-env-vars in your project. Claude Code loads it when a task matches its description.

How do I install Offline Packing Env Vars in Codex?

Run `npx skills add EvolvingLMMs-Lab/LLaVA-OneVision-2 --skill offline-packing-env-vars -a codex`. Or copy the skill folder (.opencode/skills/offline-packing-env-vars in EvolvingLMMs-Lab/LLaVA-OneVision-2) into .agents/skills/offline-packing-env-vars in your project. Codex loads it when a task matches its description.

Can I use Offline Packing Env Vars in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add EvolvingLMMs-Lab/LLaVA-OneVision-2 --skill offline-packing-env-vars -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/offline-packing-env-vars, .gemini/skills/offline-packing-env-vars, .github/skills/offline-packing-env-vars and .opencode/skills/offline-packing-env-vars in your project.

What does Offline Packing Env Vars need to run?

SKILL.md names no scripts, command-line tools or credentials: Offline Packing Env Vars is instructions for the agent only. Our summary lists: Python 3. Compatibility (from SKILL.md): opencode.

Does Offline Packing Env Vars access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Offline Packing Env Vars safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Offline Packing Env Vars use?

Offline Packing Env Vars is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Offline Packing Env Vars use?

About 3.9k tokens (SKILL.md is roughly 15k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Offline Packing Env Vars?

Skills that share tags, products or a category with Offline Packing Env Vars: Staticphp Documentation Sync (crazywhalecc/static-php-cli, 1.9k stars), Translation Diff Export (Devolutions/UniGetUI, 26k stars), Sync Translations (symfony/symfony, 31k stars) and Translation Diff Import (Devolutions/UniGetUI, 26k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Offline Packing Env Vars?

EvolvingLMMs-Lab (a GitHub organization) maintains it in EvolvingLMMs-Lab/LLaVA-OneVision-2, which has 1,216 GitHub stars. The repository holds 8 skills in this directory. The repository was last updated on October 7, 2026.

Source: EvolvingLMMs-Lab/LLaVA-OneVision-2 on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.