Agent skill

Veomni New Op

by ByteDance-Seed in ByteDance-Seed/VeOmni

A skill your agent uses when adding a new optimized kernel or operator to veomni/ops/.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Veomni New Op

skills CLI
$ npx skills add ByteDance-Seed/VeOmni --skill veomni-new-op -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ByteDance-Seed/VeOmni veomni-new-op --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ByteDance-Seed/VeOmni.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/veomni-new-op .claude/skills/veomni-new-op && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
veomni-new-op
GitHub stars
2.2k
Token cost
~3.1k tokens
SKILL.md length
1,202 words
Files
1
Skills in repo
10
Repo updated
First seen
Licence
Apache-2.0

At a glance

A skill your agent uses when adding a new optimized kernel or operator to veomni/ops/.

  • Works in 5 steps: Design → Implement → Test → …
  • Adding a new optimized kernel
  • SKILL.md covers Before You Start, VeOmni Ops Architecture, Phase 1: Design and Phase 2: Implement, plus 4 more sections
  • Calls pytest and make

What it does

Veomni New Op is an agent skill from ByteDance-Seed/VeOmni. Use this skill when adding a new optimized kernel or operator to veomni/ops/. Covers the full lifecycle: understanding VeOmni's ops architecture (KERNELREGISTRY + OpSlot dispatch, with a thin function-pointer shim for a few legacy global ops), implementing the kernel, registering it, adding tests, and documenting it. Trigger: 'add op', 'new kernel', 'add attention variant', 'new fused op', 'add triton kernel', 'optimize operator'.

Its SKILL.md is about 3.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering GPU and accelerator computing. The repository describes itself as: VeOmni: Scaling Any Modality Model Training with Model-Centric Distributed Recipe Zoo. The licence is Apache-2.0.

When your agent uses it

  • Adding a new optimized kernel
  • Operator to veomni/ops/

Example prompts

  • “add op”
  • “new kernel”
  • “add attention variant”
  • “/veomni-new-op”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Design
  2. Implement
  3. Test
  4. Document
  5. Finalize

What it can do on your machine

Read from SKILL.md and the folder at commit 8791a71. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pytest
    • make

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Veomni New Op loads about 3.1k tokens when it runs. Until then it costs about 112 tokens; SKILL.md has 1,202 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~112
When it runs · the whole SKILL.md, loaded when a task matches
~3.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from ByteDance-Seed/VeOmni at commit 8791a71, republished under its Apache-2.0 licence (© ByteDance-Seed). 1,202 words, ~3,110 tokens.

Download SKILL.mdSave it as .claude/skills/veomni-new-op/SKILL.md (or your agent's skills folder).
name
veomni-new-op
description
Use this skill when adding a new optimized kernel or operator to veomni/ops/. Covers the full lifecycle: understanding VeOmni's ops architecture (KERNEL_REGISTRY + OpSlot dispatch, with a thin function-pointer shim for a few legacy global ops), implementing the kernel, registering it, adding tests, and documenting it. Trigger: 'add op', 'new kernel', 'add attention variant', 'new fused op', 'add triton kernel', 'optimize operator'.

Before You Start

  1. Read .agents/knowledge/constraints.md — especially the "Hardware" section (NPU guards, device-agnostic helpers) and "Module-level OpSlots are shared by every model instance" under "Trainer Extensions".
  2. Read docs/design/kernel_selection.md and docs/design/unified_kernel_registry.md — understand the kernel lifecycle, the KERNEL_REGISTRY, and OpSlot dispatch.
  3. Familiarize yourself with the ops architecture below.

VeOmni Ops Architecture

Most VeOmni ops in v5 are registry-driven: a kernel registers itself in veomni.ops.kernel_registry.KERNEL_REGISTRY and is dispatched at model-build time through OpSlot instances declared in the patchgen-generated modeling files (see veomni/ops/dispatch.py and _bind_veomni_ops() in veomni/models/auto.py).

veomni/ops/
├── __init__.py          # apply_ops_patch / apply_ops_config entry points
├── kernel_registry.py   # KERNEL_REGISTRY (the single source of truth)
├── dispatch.py          # OpSlot + binding helpers
├── config/              # legacy OpSpec/BackendSpec registry: apply_global_ops()
│                        # + apply_per_model_patches() for device_patch.py models
├── kernels/             # all registry-driven kernels
│   ├── attention/       # FA2/3/4 + sequence-parallel wrappers
│   ├── cross_entropy/   # eager + liger fused CE
│   ├── deepseek_sparse_attention/
│   ├── deepseek_v4/     # TileLang sparse attention / indexer
│   ├── load_balancing_loss/
│   ├── mhc/             # TileKernels DeepSeek V4 adapters
│   ├── moe/             # fused MoE (group_gemm / quack / npu_group_gemm)
│   ├── rms_norm/        # eager / liger / batch-invariant
│   ├── rotary/          # default / triton-deterministic
│   ├── swiglu/          # eager / liger
│   └── gated_delta_rule/
├── batch_invariant_ops/ # ATen-level interception for bitwise determinism
├── liger/               # Liger kernel adapters
└── platform/            # NPU-specific helpers

Three mechanisms coexist. Pick the first one unless you have a concrete reason not to:

  1. KERNEL_REGISTRY + OpSlot (preferred for new ops). Each kernel registers itself under a (slot_name, variant) pair (e.g. ("cross_entropy_loss", "causal"), ("moe_experts", "standard")). Patchgen-generated modeling code declares matching OpSlot instances; at model-build time _bind_veomni_ops() walks the generated module, finds each OpSlot, and binds it to the concrete registry entry chosen by OpsImplementationConfig (config/registry.py).
  2. Legacy global function pointer shim (kept for a few global ops that are dispatched outside generated modeling). Public-API functions like fused_moe_forward and load_balancing_loss still expose a thin pointer that is rebound by apply_ops_config() so call sites in non-patchgen code (DeepSeek MLA inference paths, NPU custom forwards) can keep importing the public name without going through an OpSlot.
  3. Per-model device_patch.py via OpSpec/BackendSpec in ops/config/registry.py. apply_per_model_patches(hf_module, model_name, targets={op: attr}) setattr-replaces attributes on an HF module. Used by the models that have no patchgen-generated file (wan) or that need a runtime device-specific swap after generation (deepseek_v3, deepseek_v4). Those three device_patch.py files are its only callers. Do not extend this for new kernels.

Mechanism 1 covers any kernel living inside a patchgen-generated modeling file. Use 2 only when the kernel must be callable from unpatched (or non-Transformers) Python code, and 3 only when touching a model that already ships a device_patch.py.

Phase 1: Design

  1. Determine op category:

    • Registry-driven kernel (the common case, used inside patchgen-generated modeling): register under a (slot_name, variant) in KERNEL_REGISTRY and add a matching OpSlot in the relevant <model>_patch_gen_config.py. No global mutation; selection is driven by OpsImplementationConfig.
    • Global op with public API (e.g. fused_moe_forward, load_balancing_loss): expose a public function in veomni/ops/__init__.py and rebind it from apply_ops_config() based on the active OpsImplementationConfig. Only use this when a non-patchgen call site (NPU MLA forward, manual inference scripts, etc.) needs to import the kernel directly.
    • Library op (no dispatch — called directly by model code): just create the module, no registry entry needed.
    • NPU variant: add alongside the GPU implementation behind an is_torch_npu_available() guard.
  2. Decide selection mechanism: read docs/design/kernel_selection.md and docs/design/unified_kernel_registry.md to determine if you need:

    • Config field in OpsImplementationConfig (veomni/arguments/arguments_types.py)
    • Environment variable
    • Both
  3. Determine binding timing:

    • Model build time (default): registry entries are resolved by _bind_veomni_ops() in veomni/models/auto.py when a model is constructed. New kernels just need to register themselves at import time.
    • apply_ops_config() time: legacy global ops (rebound function pointers) are wired in veomni/ops/__init__.py::apply_ops_config(ops_config).

Phase 2: Implement

  1. Create the op directory under veomni/ops/kernels/<op_name>/.

  2. Implement each kernel variant in its own file (e.g. triton_kernel.py, eager.py, npu_kernel.py). Each variant declares a concrete function with the kernel's canonical signature.

  3. Register the kernel in veomni/ops/kernels/<op_name>/__init__.py. One KERNEL_REGISTRY.register(KernelSpec(...)) call per implementation — register() takes a single KernelSpec and returns None, so it is not a decorator:

    python
    from veomni.ops.kernel_registry import KERNEL_REGISTRY, HardwareRequirement, KernelSpec
    
    
    def _my_op_triton_factory():
        from .triton_kernel import my_op_triton  # imported only when selected
    
        return my_op_triton
    
    
    KERNEL_REGISTRY.register(
        KernelSpec(
            name="triton",              # impl name the user selects in the config
            op_name="my_op",            # the logical op — matches the OpSlot
            variant="standard",         # op shape, when one op has several
            factory=_my_op_triton_factory,
            hardware=HardwareRequirement(device_type="gpu"),
            description="Triton my_op",
        )
    )

    factory is a zero-argument callable returning the kernel, not the kernel itself. Keeping it lazy is what stops an optional dependency (Liger, Triton, torch_npu) from being imported just because the module was loaded. hardware is enforced at resolve() time, so an unavailable kernel fails with a clear error instead of at first use.

    Mind the two axes: (op_name, variant) identifies the slot, name identifies the implementation within it. Kernels in different variants never collide.

    Then declare a matching OpSlot in the patchgen config of every model that uses it — the arguments are (op_name, variant), not an implementation:

    python
    from veomni.ops.dispatch import OpSlot
    veomni_my_op = OpSlot("my_op", "standard")

    _bind_veomni_ops() calls slot.bind(impl_name) with the implementation selected by OpsImplementationConfig. See veomni/ops/kernels/rotary/__init__.py for a live example, and veomni/ops/README.md for the op/variant/impl table.

  4. Wire the config field (if the user needs to choose an implementation):

    • Add a field to OpsImplementationConfig in veomni/arguments/arguments_types.py.
    • Call register_op(OpSpec(name=..., config_field=..., scope=..., default=..., backends={...})) from the same veomni/ops/kernels/<op_name>/__init__.py — the mapping lives next to the kernel, not inside veomni/ops/config/registry.py, which only defines OpSpec / BackendSpec / register_op. See veomni/ops/kernels/rms_norm/__init__.py, which registers both an OpSpec and its KernelSpecs.
  5. For legacy global ops (only when needed): add the public function to veomni/ops/__init__.py and rebind it from apply_ops_config(ops_config).

  6. Async Ulysses split wrappers (only for rms_norm and rotary_pos_emb): compound Functions cannot call OpSlot. They use no-autograd (output, saved) / backward pairs in veomni/distributed/sequence_parallel/op_wrappers.py. A new backend or variant must either add a matching wrapper there, or be left off _SUPPORTED_IMPLEMENTATIONS / _SUPPORTED_VARIANTS so get_op_wrapper rejects it. KERNEL_REGISTRY coverage is not enough.

  7. NPU support:

    • Always guard NPU imports with is_torch_npu_available().
    • Put NPU implementations in a separate file (e.g., npu_kernel.py).
    • Register the NPU variant under the same slot with a distinct variant name.
Show full SKILL.md (391 more words)Show less

Phase 3: Test

  1. Add unit tests to tests/ops/. The GPU job runs this directory wholesale, so a new file needs no gpu_unit_tests.yml change. The NPU job does not — it enumerates ops files by name, so if the kernel must run on Ascend, add a line to npu_unit_tests.yml (see .agents/knowledge/testing.md):

    • Test correctness: compare output against a reference implementation (eager PyTorch)
    • Test numerical precision: verify tolerance for bf16/fp16
    • Test edge cases: empty inputs, single-element tensors, extreme shapes If the kernel only binds on SM90+, guard it so the SM89 GPU runners skip rather than fail.
  2. Add benchmark (optional but recommended for performance-critical ops):

    • Use veomni/ops/kernels/moe/_kernels/utils/benchmark_utils.py as reference
    • Compare against baseline implementation
  3. Run: pytest tests/ops/ -v

Phase 4: Document

  1. Update docs/design/kernel_selection.md:

    • Add the new op to the Quick Reference table
    • Describe the selection mechanism
  2. Update .agents/knowledge/architecture.md if the op adds a new subdirectory to veomni/ops/.

Phase 5: Finalize

  1. Run make quality.
  2. Verify the new variant shows up in KERNEL_REGISTRY.dump() and that the relevant OpSlot is rebound after build_foundation_model.
  3. Before opening the PR, run /veomni-review over the branch diff — a new kernel touches veomni/, so the gate applies.

Common Pitfalls

  • Forgetting to register in KERNEL_REGISTRY: the variant is invisible to _bind_veomni_ops() and OpSlot will fall through to its default — you'll silently exercise the wrong kernel.
  • Forgetting to add the matching OpSlot to the patchgen config: registering a kernel alone has no effect — generated modeling code must declare an OpSlot for it to be picked up.
  • Unconditional NPU imports: importing NPU modules without an is_torch_npu_available() guard crashes on GPU-only environments.
  • Binding at wrong time: registry entries are resolved when build_foundation_model runs _bind_veomni_ops(). Kernels that depend on per-model config must be picked at that point — not at module-import time.
  • New rms_norm / rotary_pos_emb backend without an async wrapper: OpSlot will bind, but async Ulysses goes through op_wrappers.py, not the registry callable. Add a split wrapper or confirm get_op_wrapper rejects the new name; do not derive the supported set from KERNEL_REGISTRY.
  • Sequence parallel interaction: ops that touch attention or loss must handle sequence parallel correctly — use get_parallel_state().sp_enabled to check and dispatch.
  • Mixed precision: fused kernels often require specific dtypes (bf16/fp16). Add assertions at the public API level to catch dtype mismatches early.
  • Not exporting public APIs: if the op provides a public function (legacy global ops), export it from veomni/ops/__init__.py's __all__.

© ByteDance-Seed, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/veomni-new-op of ByteDance-Seed/VeOmni.

Open the folder on GitHubat commit 8791a71

Compare with similar skills

Veomni New Op next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Veomni New Op compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Veomni New Op this skillByteDance-Seed/VeOmni2.2k—~3.1kAutomated safety check: PassApache-2.0
Hugging Face Local Model Evalshuggingface/skills11k2 repos~1.6kAutomated safety check: PassApache-2.0
Liger Kernel Perflinkedin/Liger-Kernel6.7k—~1.5kAutomated safety check: PassBSD-2-Clause
Hugging Face LLM Trainerhuggingface/skills11k1 repos~7.2kAutomated safety check: PassApache-2.0
MUSA GPU Training Optimizeropen-infra-skills/infra-skills141—~1.7kAutomated safety check: PassApache-2.0
Cuda Kernel OptimizerKernelFlow-ops/cuda-optimized-skill214—~4.3kAutomated safety check: PassMIT

Similar skills

  • Official

    Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.

    11k GitHub starsUsed in 2 repos~1.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Liger Kernel Perf

    linkedin/Liger-Kernel

    Optimizes the performance of existing Liger Kernel Triton kernels.

    6.7k GitHub stars~1.5k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Hugging Face LLM Trainer

    huggingface/skills

    Official

    Trains or fine-tunes language and vision models with TRL or Unsloth on Hugging Face Jobs cloud GPUs, then converts the results to GGUF.

    11k GitHub starsUsed in 1 repo~7.2k tokens
    AI & LLM EngineeringAuto-check passed
  • MUSA GPU Training Optimizer

    open-infra-skills/infra-skills

    Profiles, benchmarks and tunes AI training workloads on Moore Threads MUSA GPUs with a measurement-first process that keeps model behavior unchanged.

    141 GitHub stars~1.7k tokensUpdated 3 mo ago
    AI & LLM EngineeringAuto-check passed
  • Cuda Kernel Optimizer

    KernelFlow-ops/cuda-optimized-skill

    Iteratively optimize a CUDA/CUTLASS/Triton kernel only when strict on-device compilation, correctness, timing, and NCU evidence gates pass.

    214 GitHub stars~4.3k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Areno Debug Runtime

    inclusionAI/AReno

    Diagnose failed, hung, slow, OOM, NaN, illegal-memory-access, NCCL, compilation, rollout, or training runs in AReno.

    323 GitHub stars~486 tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed

More from ByteDance-Seed/VeOmni

All 10 skills in this repo
  • Create PR

    ByteDance-Seed/VeOmni

    Create a pull request for the current branch. An agent skill from ByteDance-Seed/VeOmni.

    2.2k GitHub stars~1.6k tokensUpdated yesterday
    Auto-check: notes
  • Veomni Debug

    ByteDance-Seed/VeOmni

    A skill your agent uses for ANY bug, error, crash, wrong output, loss divergence, gradient explosion, test failure, CUDA error, distributed training hang, checkpoint load failure, or unexpected…

    2.2k GitHub stars~2.8k tokensUpdated yesterday
    Auto-check passed
  • Veomni New Model

    ByteDance-Seed/VeOmni

    A skill your agent uses when adding support for a new model to VeOmni.

    2.2k GitHub stars~2k tokensUpdated yesterday
    Auto-check passed
  • Veomni Patchgen Model

    ByteDance-Seed/VeOmni

    Author or refresh a VeOmni model's patchgen-generated modeling under generated/ — GPU and/or NPU config, dense or MoE, text / VLM / Omni.

    2.2k GitHub stars~9.6k tokensUpdated yesterday
    Auto-check passed
  • Veomni Profile

    ByteDance-Seed/VeOmni

    A skill your agent uses for performance profiling and optimization.

    2.2k GitHub stars~1.7k tokensUpdated yesterday
    Auto-check passed
  • Veomni Review

    ByteDance-Seed/VeOmni

    Pre-PR code review gate. An agent skill from ByteDance-Seed/VeOmni.

    2.2k GitHub stars~1.9k tokensUpdated yesterday
    Auto-check passed

Questions about Veomni New Op

What does Veomni New Op do?

A skill your agent uses when adding a new optimized kernel or operator to veomni/ops/. Veomni New Op is an agent skill from ByteDance-Seed/VeOmni. Use this skill when adding a new optimized kernel or operator to veomni/ops/.

When should I use Veomni New Op?

Veomni New Op fits situations like: adding a new optimized kernel; operator to veomni/ops/.

How do I install Veomni New Op in Claude Code?

Run `npx skills add ByteDance-Seed/VeOmni --skill veomni-new-op -a claude-code`. Or copy the skill folder (.agents/skills/veomni-new-op in ByteDance-Seed/VeOmni) into .claude/skills/veomni-new-op in your project. Claude Code loads it when a task matches its description.

How do I install Veomni New Op in Codex?

Run `npx skills add ByteDance-Seed/VeOmni --skill veomni-new-op -a codex`. Or copy the skill folder (.agents/skills/veomni-new-op in ByteDance-Seed/VeOmni) into .agents/skills/veomni-new-op in your project. Codex loads it when a task matches its description.

Can I use Veomni New Op in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ByteDance-Seed/VeOmni --skill veomni-new-op -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/veomni-new-op, .gemini/skills/veomni-new-op, .github/skills/veomni-new-op and .opencode/skills/veomni-new-op in your project.

What does Veomni New Op need to run?

Going by SKILL.md and its folder, Veomni New Op needs the command-line tools its instructions call (pytest and make). Our summary lists: Python 3.

Does Veomni New Op access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Veomni New Op safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Veomni New Op use?

Veomni New Op is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Veomni New Op use?

About 3.1k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Veomni New Op?

Skills that share tags, products or a category with Veomni New Op: Hugging Face Local Model Evals (huggingface/skills, 11k stars), Liger Kernel Perf (linkedin/Liger-Kernel, 6.7k stars), Hugging Face LLM Trainer (huggingface/skills, 11k stars) and MUSA GPU Training Optimizer (open-infra-skills/infra-skills, 141 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Veomni New Op?

ByteDance-Seed (a GitHub organization) maintains it in ByteDance-Seed/VeOmni, which has 2,235 GitHub stars. The repository holds 10 skills in this directory. The repository was last updated on October 10, 2026.

Source: ByteDance-Seed/VeOmni on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.