Agent skill

Cu Lengths Attention Flow

by EvolvingLMMs-Lab in EvolvingLMMs-Lab/LLaVA-OneVision-2

Bilingual guide for understanding how culengths controls attention behavior across ViT and LLM stages, and how patchpositions scope differs between the two

Apache-2.0Auto-check passedWriting & Content

Install Cu Lengths Attention Flow

skills CLI
$ npx skills add EvolvingLMMs-Lab/LLaVA-OneVision-2 --skill cu-lengths-attention-flow -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install EvolvingLMMs-Lab/LLaVA-OneVision-2 cu-lengths-attention-flow --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/EvolvingLMMs-Lab/LLaVA-OneVision-2.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.opencode/skills/cu-lengths-attention-flow .claude/skills/cu-lengths-attention-flow && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
cu-lengths-attention-flow
GitHub stars
1.2k
Token cost
~3.1k tokens
SKILL.md length
964 words
Files
1
Skills in repo
8
Repo updated
First seen
Licence
Apache-2.0

At a glance

Bilingual guide for understanding how culengths controls attention behavior across ViT and LLM stages, and how patchpositions scope differs between the two

  • Works in 4 steps: cu_lengths Generation / cu_lengths 的产生 → Forward Function Branching / 前向函数分支 → LLM Attention Behavior / LLM 层的… → …
  • Tasks that involve Translation
  • SKILL.md covers Purpose / 用途, Key Files / 关键文件, Core Concept: Two Independent… and Mechanism Details / 机制详解, plus 4 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Cu Lengths Attention Flow is an agent skill from EvolvingLMMs-Lab/LLaVA-OneVision-2. Bilingual guide for understanding how culengths controls attention behavior across ViT and LLM stages, and how patchpositions scope differs between the two

Its SKILL.md is about 3.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts. Compatibility notes: opencode

It sits in Writing & Content, covering Translation. The repository describes itself as: Fully Open Framework for Democratized Multimodal Training. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Translation

Example prompts

  • “/cu-lengths-attention-flow”

Requirements

  • Python 3
  • Compatibility (from SKILL.md): opencode

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. cu_lengths Generation / cu_lengths 的产生
  2. Forward Function Branching / 前向函数分支
  3. LLM Attention Behavior / LLM 层的 Attention 行为
  4. Sequence Parallelism Padding / 序列并行填充

What it can do on your machine

Read from SKILL.md and the folder at commit 6ef16b1. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    opencode

    From compatibility in the SKILL.md frontmatter.

Context cost

Cu Lengths Attention Flow loads about 3.1k tokens when it runs. Until then it costs about 46 tokens; SKILL.md has 964 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~46
When it runs · the whole SKILL.md, loaded when a task matches
~3.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from EvolvingLMMs-Lab/LLaVA-OneVision-2 at commit 6ef16b1, republished under its Apache-2.0 licence (© EvolvingLMMs-Lab). 964 words, ~3,076 tokens.

Download SKILL.mdSave it as .claude/skills/cu-lengths-attention-flow/SKILL.md (or your agent's skills folder).
name
cu-lengths-attention-flow
description
Bilingual guide for understanding how cu_lengths controls attention behavior across ViT and LLM stages, and how patch_positions scope differs between the two
compatibility
opencode
metadata.domain
model-architecture
metadata.framework
megatron-energon
metadata.repo
llava-onevision2

Purpose / 用途

Use this skill when reasoning about attention boundaries in the LLaVA-OneVision2 forward pass — specifically how cu_lengths and patch_positions control attention at different stages of the model.

在分析 LLaVA-OneVision2 前向传播中的 attention 边界时使用这个 skill——具体来说,cu_lengths 和 patch_positions 如何在模型的不同阶段控制 attention。

This skill is specifically for:

  • Understanding the difference between ViT-level and LLM-level attention control
  • Debugging packed vs non-packed attention behavior
  • Reasoning about cross-sample isolation in packed sequences
  • Understanding why patch_positions grouping does NOT carry into the LLM

这个 skill 专门用于:

  • 理解 ViT 层和 LLM 层 attention 控制的区别
  • 调试 packed 和 non-packed 的 attention 行为
  • 分析 packed 序列中跨样本隔离机制
  • 理解为什么 patch_positions 的分组不会延续到 LLM 中

Key Files / 关键文件

FileRole
aiak_training_llm/train/pretrain/pretrain_llava_onevision2.pyForward function — decides packed vs non-packed path based on cu_lengths shape
aiak_training_llm/train/sft/utils.py_get_packed_sequence_params() — builds PackedSeqParams from attention_mask for SFT
aiak_training_llm/data/multimodal/task_encoder.pybatch() — sets cu_lengths to [[0]] (dummy) for non-packed, or stacks real cu_lengths for packed
aiak_training_llm/data/multimodal/task_encoder.pypack_selected_samples() — constructs cu_lengths = [0, len_1, len_1+len_2, ...] for offline packed data
aiak_training_llm/models/llava_onevision2/onevision_encoder_model.pyViT encoder — uses patch_positions for local/shared attention
aiak_training_llm/data/multimodal/qwen2vl_task_encoder.pyprocess_sft_qa() — generates patch_positions from image_grid_thw

Core Concept: Two Independent Attention Control Mechanisms / 核心概念:两套独立的 Attention 控制机制

Overview Diagram / 概览图
┌─────────────────────────────────────────────────┐
│  ViT Encoder                                    │
│                                                 │
│  Control: patch_positions (temporal dimension)  │
│  Effect:  Local/shared attention                │
│           e.g. 4 images share one attention     │
│           window via same temporal index        │
│                                                 │
│  Output:  visual embeddings                     │
└──────────────────┬──────────────────────────────┘
                   │  (embeddings replace image
                   │   placeholder tokens)
                   ▼
┌─────────────────────────────────────────────────┐
│  LLM Decoder                                    │
│                                                 │
│  Control: cu_lengths (cumulative sub-seq lens)  │
│  Effect:  Determines attention domain           │
│                                                 │
│  NON-PACKED: cu_lengths == [[0]]                │
│    → full causal attention                      │
│    → ALL tokens see ALL previous tokens         │
│    → patch_positions grouping is GONE           │
│                                                 │
│  PACKED: cu_lengths = [0, a, a+b, ...]          │
│    → block-diagonal causal attention            │
│    → sub-sequences isolated from each other     │
│    → within each sub-seq: full causal           │
└─────────────────────────────────────────────────┘

Mechanism Details / 机制详解

1. cu_lengths Generation / cu_lengths 的产生
Source A: Offline Packed Data (PackedCaptioningSample)
python
# task_encoder.py → pack_selected_samples()
cu_lengths = [0]
for sample in samples:
    current_length += sample.total_len
    cu_lengths.append(current_length)
# Result: [0, 512, 1024, 1389] — 3 sub-samples packed together

Each sub-sample was independently encoded (tokenized + image processed) then concatenated into one long sequence. cu_lengths records the boundaries.

每个子样本独立编码(tokenize + 图像处理),然后拼接成一个长序列。cu_lengths 记录边界。

Source B: Non-Packed Single Samples
python
# task_encoder.py → batch()
if self.is_packing_enabled or int(os.environ.get("OFFLINE_PACKED_DATA", 0)) == 1:
    cu_lengths = torch.stack([s.cu_lengths for s in samples])
else:
    cu_lengths = torch.tensor([[0]], dtype=torch.int32)  # dummy value

Non-packed samples get cu_lengths = [[0]] with shape [1, 1].

非 packed 样本得到 cu_lengths = [[0]],shape 为 [1, 1]。

2. Forward Function Branching / 前向函数分支

In pretrain_llava_onevision2.py:

python
if cu_lengths.shape == torch.Size([1, 1]):
    # ===== NON-PACKED PATH =====
    # Uses attn_mask for padding_causal attention
    # Every token attends to all previous tokens (full causal)
    # packed_seq_params = None
    for i in range(attn_mask.shape[0]):
        loss_mask[i, (attn_mask[i] == False).sum() - 1] = 0
else:
    # ===== PACKED PATH =====
    # micro-batch-size must be 1 for packing
    assert cu_lengths.shape[0] == 1
    attn_mask = None  # not needed — cu_seqlens defines boundaries
    packed_seq_params = PackedSeqParams(
        qkv_format="thd",
        cu_seqlens_q=cu_lengths[0],      # → Flash Attention kernel
        cu_seqlens_kv=cu_lengths[0],
        max_seqlen_q=max_lengths[0].item(),
        max_seqlen_kv=max_lengths[0].item(),
    )
3. LLM Attention Behavior / LLM 层的 Attention 行为

Non-packed (cu_lengths == [[0]]) → Full Causal Attention:

  • The entire sequence is ONE attention domain
  • Token at position i can attend to positions 0..i
  • ALL visual tokens (from any image) + ALL text tokens are mutually visible
  • Compute cost: O(n²) where n = total sequence length
  • packed_seq_params = None → standard causal mask

整个序列是一个 attention 域,所有 visual token(来自任何图片)和所有 text token 互通可见。计算量 O(n²)。

Packed (cu_lengths = [0, a, a+b, ...]) → Block-Diagonal Causal Attention:

  • Each sub-sequence [cu_lengths[i], cu_lengths[i+1]) is an independent attention domain
  • Within each sub-sequence: full causal attention
  • Between sub-sequences: ZERO attention (completely isolated)
  • Implemented via Flash Attention's cu_seqlens parameter
  • Compute cost: O(a² + b² + c² + ...) << O((a+b+c+...)²)

每个子序列 [cu_lengths[i], cu_lengths[i+1]) 是独立的 attention 域,子序列之间完全隔离。通过 Flash Attention 的 cu_seqlens 参数实现。

4. Sequence Parallelism Padding / 序列并行填充

When args.sequence_parallel is enabled, the sequence must be divisible by TP size (and TP×CP×2 if CP > 1). For packed sequences, padding tokens are appended as a dummy extra sub-sequence:

当启用 sequence_parallel 时,序列长度必须被 TP size 整除。对 packed 序列,padding token 作为一个额外的 dummy 子序列追加:

python
if packed_seq_params is not None:
    new_end_q = packed_seq_params.cu_seqlens_q[-1:] + pad_size
    packed_seq_params = PackedSeqParams(
        cu_seqlens_q=torch.cat([packed_seq_params.cu_seqlens_q, new_end_q]),
        cu_seqlens_kv=torch.cat([packed_seq_params.cu_seqlens_kv, new_end_kv]),
        max_seqlen_q=max(packed_seq_params.max_seqlen_q, pad_size),
        max_seqlen_kv=max(packed_seq_params.max_seqlen_kv, pad_size),
    )

Critical Insight: patch_positions Scope / 关键洞察:patch_positions 的作用域

ViT Layer: patch_positions Controls Attention Grouping

In the ViT encoder, patch_positions has a temporal dimension (t, h, w). Images sharing the same temporal index share one attention window. For example, 4 images treated as "video frames" share attention via their temporal coordinates.

在 ViT encoder 中,patch_positions 有时间维度 (t, h, w)。共享相同时间索引的图片共享一个 attention window。例如,4 张图片被当作"视频帧"通过时间坐标共享 attention。

LLM Layer: patch_positions Has NO Effect on Attention

patch_positions is NOT used to control LLM attention. It is only passed through the data pipeline for potential use in position embeddings or other purposes, but the LLM's attention boundaries are controlled EXCLUSIVELY by cu_lengths.

patch_positions 不控制 LLM 的 attention。 它只是在数据 pipeline 中传递,可能用于 position embedding 等目的,但 LLM 的 attention 边界完全由 cu_lengths 控制。

This means:

这意味着:

ScenarioViT AttentionLLM Attention
4 images with shared patch_positions temporal index, non-packed4 images share attention windowALL tokens (all 4 images + text) in full causal — no grouping
4 images with shared patch_positions temporal index, packed (separate sub-samples)4 images share attention windowEach sub-sample isolated via cu_seqlens — inter-sample isolation
Single image, non-packedStandard ViT attentionFull causal over entire sequence
Show full SKILL.md (375 more words)Show less
Why This Matters / 为什么这很重要

If you assume ViT-level grouping persists into the LLM, you will misunderstand the compute profile:

如果你假设 ViT 层的分组延续到 LLM 中,会误解计算特征:

  • Non-packed: LLM always does full causal attention over the entire sequence. A sample with 4 high-res images has O((4×img_tokens + text_tokens)²) attention cost — there is NO per-image isolation in the LLM.

  • Packed: The isolation is between SAMPLES (sub-sequences), not between images within a sample. A packed sample containing samples A (2 images) and B (1 image) isolates A from B, but within A, both images + text are fully visible to each other.

  • Non-packed: LLM 总是对整个序列做 full causal attention。一个包含 4 张高分辨率图片的样本,attention 代价为 O((4×img_tokens + text_tokens)²)——LLM 中没有按图片隔离。

  • Packed: 隔离是在样本(子序列)之间,不是在同一样本内的图片之间。一个包含样本 A(2 张图)和样本 B(1 张图)的 packed 样本,A 和 B 互相隔离,但 A 内部的两张图 + 文本完全互通。


SFT Path: Attention Mask Based cu_seqlens / SFT 路径:基于 attention_mask 的 cu_seqlens

In the SFT training path (sft/utils.py), cu_seqlens can also be derived from the attention mask using sample-ID encoding:

在 SFT 训练路径中,cu_seqlens 也可以从 attention mask 推导:

python
# sft/utils.py → _get_packed_sequence_params()
# attention_mask encodes sample IDs: [[1,1,2,2,2,3,3,4,5,5,5,0,0]]
# → cu_seqlens = [0, 2, 5, 7, 8, 11, 13]
reduced_mask = torch.bincount(attention_mask.view(-1), minlength=max_num + 1)
cu_seqlens = reduced_mask[1:].cumsum(dim=0).to(torch.int32)
cu_seqlens[-1] = attention_mask.shape[1]  # include padding
cu_seqlens = torch.cat((zero, cu_seqlens))

This achieves the same block-diagonal attention as the pretrain path's cu_lengths mechanism.

这与 pretrain 路径的 cu_lengths 机制实现相同的 block-diagonal attention。


Quick Reference Table / 快速参考表

FieldWhere SetWhat It Controls
cu_lengthstask_encoder.batch() or pack_selected_samples()LLM attention boundaries (packed vs full causal)
packed_seq_paramspretrain_*.py forward functionFlash Attention kernel parameter (cu_seqlens_q/kv)
patch_positionsqwen2vl_task_encoder.process_sft_qa()ViT local attention grouping (temporal dimension)
attn_maskencode_sample()Padding mask for non-packed; set to None for packed
max_lengthspack_selected_samples()Max sub-sequence length in packed sample (for Flash Attention)
Shape of cu_lengthsMeaningLLM Attention Type
[1, 1] (value [[0]])Non-packed / dummyFull causal
[1, P] where P > 1Packed with P-1 sub-samplesBlock-diagonal causal

Common Pitfalls / 常见误区

  1. Assuming ViT attention grouping carries into LLM — It does NOT. patch_positions only affects ViT; LLM uses cu_lengths.

    假设 ViT 的 attention 分组延续到 LLM — 不会。patch_positions 只影响 ViT;LLM 使用 cu_lengths。

  2. Confusing packed sample isolation with image-level isolation — cu_lengths boundaries separate SAMPLES, not images within a sample.

    混淆 packed 样本隔离和图片级隔离 — cu_lengths 边界分隔的是样本,不是样本内的图片。

  3. Forgetting micro-batch-size=1 constraint for packing — The code asserts cu_lengths.shape[0] == 1 in the packed path.

    忘记 packing 要求 micro-batch-size=1 — 代码在 packed 路径中断言 cu_lengths.shape[0] == 1。

  4. Ignoring SP padding for packed sequences — When sequence parallelism is enabled, padding tokens are added as a dummy sub-sequence in cu_seqlens, not ignored.

    忽略 packed 序列的 SP 填充 — 启用序列并行时,padding token 作为 dummy 子序列加入 cu_seqlens,不是被忽略。

© EvolvingLMMs-Lab, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .opencode/skills/cu-lengths-attention-flow of EvolvingLMMs-Lab/LLaVA-OneVision-2.

Open the folder on GitHubat commit 6ef16b1

Compare with similar skills

Cu Lengths Attention Flow next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Cu Lengths Attention Flow compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Cu Lengths Attention Flow this skillEvolvingLMMs-Lab/LLaVA-OneVision-21.2k—~3.1kAutomated safety check: PassApache-2.0
Translation Diff ExportDevolutions/UniGetUI26k—~1.1kAutomated safety check: PassMIT
Sync Translationssymfony/symfony31k—~1.9kAutomated safety check: PassMIT
Translation Diff ImportDevolutions/UniGetUI26k—~750Automated safety check: PassMIT
Translation Diff TranslateDevolutions/UniGetUI26k—~934Automated safety check: PassMIT
Generate Translationspayloadcms/payload45k—~1.1kAutomated safety check: PassMIT

Similar skills

  • Translation Diff Export

    Devolutions/UniGetUI

    Compares UniGetUI JSON locale files against English, identifies untranslated or source-changed keys, and generates patch, reference, and handoff files for a target language.

    26k GitHub stars~1.1k tokensUpdated yesterday
    Writing & ContentAuto-check passed
  • Sync Translations

    symfony/symfony

    Synchronize translation catalogs across maintained Symfony branches: find messages that newer branches added to the English catalogs but that are still missing from the oldest maintained branch…

    31k GitHub stars~1.9k tokensUpdated yesterday
    Writing & ContentAuto-check passed
  • Translation Diff Import

    Devolutions/UniGetUI

    Merges translated key-value pairs from a UniGetUI JSON localization patch back into the full language file and validates the merged result.

    26k GitHub stars~750 tokensUpdated yesterday
    Writing & ContentAuto-check passed
  • Translation Diff Translate

    Devolutions/UniGetUI

    Translates a sparse UniGetUI JSON language patch, writes completed entries into the working copy, preserves placeholders and terminology, and prepares the patch for merge-back.

    26k GitHub stars~934 tokensUpdated yesterday
    Writing & ContentAuto-check passed
  • Generate Translations

    payloadcms/payload

    A skill your agent uses when new translation keys are added to packages to generate new translations strings

    45k GitHub stars~1.1k tokensUpdated today
    Writing & ContentAuto-check passed
  • Drives long-form fiction, scripts, storyboards, interactive films and long-document translation through InkOS, with every change made by a typed action.

    10k GitHub starsUsed in 1 repo~1.1k tokens
    Writing & ContentAuto-check passed

More from EvolvingLMMs-Lab/LLaVA-OneVision-2

All 8 skills in this repo
  • Commit Message

    EvolvingLMMs-Lab/LLaVA-OneVision-2

    Guide for writing clear, consistent git commit messages following this repository's conventions

    1.2k GitHub stars~1.9k tokensUpdated yesterday
    Auto-check passed
  • Distributed Offline Packing

    EvolvingLMMs-Lab/LLaVA-OneVision-2

    Bilingual guide for running offlinepacking/autopipe.sh across multiple nodes to produce padding-free packed WebDataset shards for SFT, with Energon Metadataset assembly

    1.2k GitHub stars~2.6k tokensUpdated yesterday
    Auto-check passed
  • Length Pool Sort Dataset

    EvolvingLMMs-Lab/LLaVA-OneVision-2

    Bilingual guide for understanding LengthPoolSortDataset cross-rank length synchronization mechanism in multi-GPU training

    1.2k GitHub stars~2k tokensUpdated yesterday
    Auto-check passed
  • Llava Onevision2 Consistency

    EvolvingLMMs-Lab/LLaVA-OneVision-2

    Bilingual guide for running and interpreting LLaVA-OneVision2 HF vs Megatron consistency checks across TP and PP settings

    1.2k GitHub stars~4.1k tokensUpdated yesterday
    Auto-check passed
  • Offline Packing Env Vars

    EvolvingLMMs-Lab/LLaVA-OneVision-2

    Bilingual guide for the OFFLINEPACKINGBMR and OFFLINEPACKEDDATA environment variables that control LLaVA-OneVision2 training-side packing — what each gate does, why both must be enabled together…

    1.2k GitHub stars~3.9k tokensUpdated yesterday
    Auto-check passed
  • Merge Ov2

    EvolvingLMMs-Lab/LLaVA-OneVision-2

    Bilingual guide for merging ViT + LLM into LlavaOnevision2 HF checkpoint and validating weight/inference consistency

    1.2k GitHub stars~7.6k tokensUpdated yesterday
    Auto-check passed

Questions about Cu Lengths Attention Flow

What does Cu Lengths Attention Flow do?

Bilingual guide for understanding how culengths controls attention behavior across ViT and LLM stages, and how patchpositions scope differs between the two. Cu Lengths Attention Flow is an agent skill from EvolvingLMMs-Lab/LLaVA-OneVision-2.

When should I use Cu Lengths Attention Flow?

Cu Lengths Attention Flow fits situations like: tasks that involve Translation.

How do I install Cu Lengths Attention Flow in Claude Code?

Run `npx skills add EvolvingLMMs-Lab/LLaVA-OneVision-2 --skill cu-lengths-attention-flow -a claude-code`. Or copy the skill folder (.opencode/skills/cu-lengths-attention-flow in EvolvingLMMs-Lab/LLaVA-OneVision-2) into .claude/skills/cu-lengths-attention-flow in your project. Claude Code loads it when a task matches its description.

How do I install Cu Lengths Attention Flow in Codex?

Run `npx skills add EvolvingLMMs-Lab/LLaVA-OneVision-2 --skill cu-lengths-attention-flow -a codex`. Or copy the skill folder (.opencode/skills/cu-lengths-attention-flow in EvolvingLMMs-Lab/LLaVA-OneVision-2) into .agents/skills/cu-lengths-attention-flow in your project. Codex loads it when a task matches its description.

Can I use Cu Lengths Attention Flow in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add EvolvingLMMs-Lab/LLaVA-OneVision-2 --skill cu-lengths-attention-flow -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/cu-lengths-attention-flow, .gemini/skills/cu-lengths-attention-flow, .github/skills/cu-lengths-attention-flow and .opencode/skills/cu-lengths-attention-flow in your project.

What does Cu Lengths Attention Flow need to run?

SKILL.md names no scripts, command-line tools or credentials: Cu Lengths Attention Flow is instructions for the agent only. Our summary lists: Python 3. Compatibility (from SKILL.md): opencode.

Does Cu Lengths Attention Flow access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Cu Lengths Attention Flow safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Cu Lengths Attention Flow use?

Cu Lengths Attention Flow is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Cu Lengths Attention Flow use?

About 3.1k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Cu Lengths Attention Flow?

Skills that share tags, products or a category with Cu Lengths Attention Flow: Translation Diff Export (Devolutions/UniGetUI, 26k stars), Sync Translations (symfony/symfony, 31k stars), Translation Diff Import (Devolutions/UniGetUI, 26k stars) and Translation Diff Translate (Devolutions/UniGetUI, 26k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Cu Lengths Attention Flow?

EvolvingLMMs-Lab (a GitHub organization) maintains it in EvolvingLMMs-Lab/LLaVA-OneVision-2, which has 1,216 GitHub stars. The repository holds 8 skills in this directory. The repository was last updated on October 7, 2026.

Source: EvolvingLMMs-Lab/LLaVA-OneVision-2 on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.