Agent skill

Veomni Profile

by ByteDance-Seed in ByteDance-Seed/VeOmni

A skill your agent uses for performance profiling and optimization.

Apache-2.0Auto-check passedDevelopment

Install Veomni Profile

skills CLI
$ npx skills add ByteDance-Seed/VeOmni --skill veomni-profile -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ByteDance-Seed/VeOmni veomni-profile --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ByteDance-Seed/VeOmni.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/veomni-profile .claude/skills/veomni-profile && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
veomni-profile
GitHub stars
2.2k
Token cost
~1.7k tokens
SKILL.md length
567 words
Files
1
Skills in repo
10
Repo updated
First seen
Licence
Apache-2.0

At a glance

A skill your agent uses for performance profiling and optimization.

  • Works in 5 steps: Configure Profiling → Run Training → Collect and Analyze → …
  • Performance profiling and optimization
  • SKILL.md covers VeOmni Profiling Infrastructure, Mode 1: Analyze Existing…, Mode 2: Generate Profiles… and NPU (Ascend) Profiling
  • Calls python

What it does

Veomni Profile is an agent skill from ByteDance-Seed/VeOmni. Use this skill for performance profiling and optimization. Two modes: (1) Analyze existing profile files (Chrome traces, memory snapshots) — write scripts to parse and summarize metrics per user requirements. (2) Generate profiles during development — configure ProfileConfig, run training, collect traces, analyze bottlenecks, and suggest optimizations. Trigger: 'profile', 'performance', 'slow', 'MFU', 'throughput', 'bottleneck', 'memory usage', 'trace', 'optimize training speed'.

Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Development, covering Performance optimization. It works with CUDA. The repository describes itself as: VeOmni: Scaling Any Modality Model Training with Model-Centric Distributed Recipe Zoo. The licence is Apache-2.0.

When your agent uses it

  • Performance profiling and optimization
  • Tasks that involve Performance optimization

Example prompts

  • “profile”
  • “performance”
  • “throughput”
  • “/veomni-profile”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Configure Profiling
  2. Run Training
  3. Collect and Analyze
  4. Optimize
  5. Validate

What it can do on your machine

Read from SKILL.md and the folder at commit 8791a71. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Veomni Profile loads about 1.7k tokens when it runs. Until then it costs about 125 tokens; SKILL.md has 567 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~125
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from ByteDance-Seed/VeOmni at commit 8791a71, republished under its Apache-2.0 licence (© ByteDance-Seed). 567 words, ~1,743 tokens.

Download SKILL.mdSave it as .claude/skills/veomni-profile/SKILL.md (or your agent's skills folder).
name
veomni-profile
description
Use this skill for performance profiling and optimization. Two modes: (1) Analyze existing profile files (Chrome traces, memory snapshots) — write scripts to parse and summarize metrics per user requirements. (2) Generate profiles during development — configure ProfileConfig, run training, collect traces, analyze bottlenecks, and suggest optimizations. Trigger: 'profile', 'performance', 'slow', 'MFU', 'throughput', 'bottleneck', 'memory usage', 'trace', 'optimize training speed'.

VeOmni Profiling Infrastructure

Key components:

ComponentLocationPurpose
ProfileConfigveomni/arguments/arguments_types.pyConfig fields: enable, start_step, end_step, trace_dir, profile_memory, with_stack, etc.
create_profiler()veomni/utils/helper.pyBuilds torch.profiler.profile (CUDA) or torch_npu.profiler (NPU) with schedule
ProfileTraceCallbackveomni/trainer/callbacks/trace_callback.pyIntegrates profiler into the training loop via BaseTrainer
VeomniFlopsCounterveomni/utils/count_flops.pyAnalytical FLOPs/MFU computation per model family
EnvironMeterveomni/utils/helper.pyStep-level throughput metrics (tokens/s, FLOPs, MFU)
merge_chrome_trace.pyscripts/profile/merge_chrome_trace.pyMerge multi-rank Chrome traces for unified viewing

Output formats:

  • Chrome trace: veomni_rank{R}_{timestamp}.pt.trace.json.gz — viewable in chrome://tracing or Perfetto
  • Memory snapshot: .pkl file via torch.cuda.memory._dump_snapshot — viewable with PyTorch Memory Viz

Mode 1: Analyze Existing Profile Files

User provides one or more profile files (Chrome traces, memory snapshots, logs). Write scripts to parse and analyze them.

Steps
  1. Identify file types: .json.gz / .json (Chrome trace), .pkl (memory snapshot), .log / .txt (training logs with throughput metrics).

  2. Understand the analysis goal — ask the user what they want to know:

    • Kernel-level breakdown (which CUDA kernels dominate wall time?)
    • Communication vs computation ratio (NCCL all-reduce, all-to-all, all-gather time)
    • Memory high-water mark and allocation timeline
    • Per-step time breakdown (forward, backward, optimizer, data loading)
    • MFU / hardware utilization
    • Comparison across multiple profiles (e.g. before/after optimization, different parallelism configs)
  3. Write an analysis script using torch.profiler APIs or raw JSON parsing:

    python
    import json, gzip
    from collections import defaultdict
    
    def load_chrome_trace(path):
        opener = gzip.open if path.endswith('.gz') else open
        with opener(path, 'rt') as f:
            return json.load(f)
    
    def analyze_kernel_time(trace):
        """Group events by kernel name, sum durations."""
        kernel_times = defaultdict(float)
        for event in trace.get('traceEvents', []):
            if event.get('cat') == 'kernel':
                kernel_times[event['name']] += event.get('dur', 0)
        return sorted(kernel_times.items(), key=lambda x: -x[1])

    Adapt the script to the user's specific analysis goal. Output tables, summaries, or CSV for further processing.

  4. For multi-rank traces: use scripts/profile/merge_chrome_trace.py to merge before analysis, or analyze per-rank and compare.

  5. For memory snapshots: load with pickle, analyze allocation records, identify peak usage and largest tensors.

  6. Present findings: summarize top bottlenecks, compute/comm ratio, and actionable optimization suggestions.


Mode 2: Generate Profiles During Development

Actively profile a training run to identify performance bottlenecks or validate optimizations.

Step 1: Configure Profiling

Add or modify the profile section in the training YAML config:

yaml
train:
  profile:
    enable: true
    start_step: 5        # skip warmup steps
    end_step: 10         # capture 5 steps
    trace_dir: ./profile_output
    record_shapes: true
    profile_memory: true  # enable memory snapshot (CUDA only)
    with_stack: true      # capture Python call stacks
    with_modules: true    # annotate with nn.Module names
    rank0_only: true      # profile only rank 0 to reduce overhead

Or pass via CLI overrides: --train.profile.enable=true --train.profile.start_step=5 ...

Step 2: Run Training
bash
source .venv/bin/activate
# Single GPU
python tasks/train_text.py --config configs/text/<model>.yaml

# Multi-GPU (profile will capture per-rank traces)
torchrun --nproc_per_node=8 tasks/train_text.py --config configs/text/<model>.yaml
Show full SKILL.md (265 more words)Show less
Step 3: Collect and Analyze
  1. Locate outputs in trace_dir:

    • veomni_rank*_.pt.trace.json.gz — Chrome trace
    • veomni_rank*_.pkl — memory snapshot (if profile_memory: true)
  2. Write analysis scripts as in Mode 1 to extract the metrics the user needs.

  3. Quick analysis shortcuts:

    • Kernel time breakdown: parse Chrome trace events with cat == 'kernel'
    • NCCL communication: filter events with names matching nccl (e.g. ncclAllReduceRingLLKernel)
    • Forward/backward split: use with_modules trace annotations to separate phases
    • Memory peak: load .pkl snapshot, find max allocated_bytes
    • MFU from logs: EnvironMeter already logs flops_achieved and flops_promised — grep training logs
  4. For multi-rank comparison: merge traces with scripts/profile/merge_chrome_trace.py or analyze per-rank to find stragglers.

Step 4: Optimize

Based on findings, suggest and implement optimizations:

BottleneckTypical solutions
Attention kernels dominateSwitch to FlashAttention 3/4 (veomni/ops/kernels/attention/), check FA is actually active
NCCL communication > 30%Increase compute/comm overlap, adjust FSDP reshard policy, try async SP
Memory OOM / high peakEnable activation checkpointing, reduce micro-batch size, check for memory leaks
Data loading stallsIncrease num_workers, enable prefetch, check I/O throughput
Low MFU (< 40%)Check dtype (bf16 vs fp32), verify tensor cores are used, check for host-device syncs
Uneven per-rank timeCheck MoE load balancing, verify data distribution across ranks
Step 5: Validate

After optimization:

  1. Re-profile with the same config to compare before/after.
  2. Verify training correctness is preserved (loss matches baseline).
  3. Document the optimization and results.

NPU (Ascend) Profiling

On NPU, create_profiler() uses torch_npu.profiler instead of torch.profiler. Key differences:

  • Output format includes AiC (Ascend insight Counters) metrics.
  • Memory profiling uses NPU-specific APIs.
  • Analysis tools differ — use Ascend Insight instead of Chrome tracing.
  • Always guard NPU-specific analysis code with is_torch_npu_available().

© ByteDance-Seed, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/veomni-profile of ByteDance-Seed/VeOmni.

Open the folder on GitHubat commit 8791a71

Compare with similar skills

Veomni Profile next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Veomni Profile compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Veomni Profile this skillByteDance-Seed/VeOmni2.2k—~1.7kAutomated safety check: PassApache-2.0
The Art of Debuggingstas00/the-art-of-debugging1.7k—~6.1kAutomated safety check: NotesCC-BY-SA-4.0
Cudatechnillogue/ptx-isa-markdown229—~2.5kAutomated safety check: PassNone
Torch Profiler Layer TrackBBuf/AI-Infra-Auto-Driven-SKILLS938—~2kAutomated safety check: PassNone
Cppcrazyguitar/cppcheatsheet290—~1.8kAutomated safety check: PassMIT
LLM Torch Profiler Trace AnalysisBBuf/AI-Infra-Auto-Driven-SKILLS938—~2.8kAutomated safety check: PassNone

Similar skills

  • The Art of Debugging

    stas00/the-art-of-debugging

    Condensed debugging method and tool recipes for Unix, Python and PyTorch programs: crashes, hangs, segfaults, wrong output, CUDA OOM, NaN values and slowness.

    1.7k GitHub stars~6.1k tokensUpdated 4 days ago
    DevelopmentAuto-check: notes
  • Cuda

    technillogue/ptx-isa-markdown

    CUDA kernel development, debugging, and performance optimization for Claude Code.

    229 GitHub stars~2.5k tokensUpdated 9 mo ago
    DevelopmentAuto-check passed
  • Torch Profiler Layer Track

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Adds verified layer guides such as L0 and L1 and compact GPU lanes to an existing Torch Profiler Chrome trace, changing how it looks but not how it ran.

    938 GitHub stars~2k tokensUpdated 6 days ago
    DevelopmentAuto-check passed
  • Cpp

    crazyguitar/cppcheatsheet

    Comprehensive C/C++ programming reference covering everything from C11-C23 and C++11-C++23, system programming, CUDA GPU computing, debugging tools, Rust interop, and advanced topics.

    290 GitHub stars~1.8k tokensUpdated yesterday
    DevelopmentAuto-check passed
  • LLM Torch Profiler Trace Analysis

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.

    938 GitHub stars~2.8k tokensUpdated 6 days ago
    AI & LLM EngineeringAuto-check passed
  • Tilelang Skill

    slowlyC/agent-gpu-skills

    Write, debug, and optimize TileLang kernels from local upstream language, JIT, autotuning, profiling, compiler, test, and example source.

    169 GitHub stars~1.8k tokensUpdated 2 mo ago
    DevelopmentAuto-check passed

More from ByteDance-Seed/VeOmni

All 10 skills in this repo
  • Create PR

    ByteDance-Seed/VeOmni

    Create a pull request for the current branch. An agent skill from ByteDance-Seed/VeOmni.

    2.2k GitHub stars~1.6k tokensUpdated yesterday
    Auto-check: notes
  • Veomni Debug

    ByteDance-Seed/VeOmni

    A skill your agent uses for ANY bug, error, crash, wrong output, loss divergence, gradient explosion, test failure, CUDA error, distributed training hang, checkpoint load failure, or unexpected…

    2.2k GitHub stars~2.8k tokensUpdated yesterday
    Auto-check passed
  • Veomni New Model

    ByteDance-Seed/VeOmni

    A skill your agent uses when adding support for a new model to VeOmni.

    2.2k GitHub stars~2k tokensUpdated yesterday
    Auto-check passed
  • Veomni New Op

    ByteDance-Seed/VeOmni

    A skill your agent uses when adding a new optimized kernel or operator to veomni/ops/.

    2.2k GitHub stars~3.1k tokensUpdated yesterday
    Auto-check passed
  • Veomni Patchgen Model

    ByteDance-Seed/VeOmni

    Author or refresh a VeOmni model's patchgen-generated modeling under generated/ — GPU and/or NPU config, dense or MoE, text / VLM / Omni.

    2.2k GitHub stars~9.6k tokensUpdated yesterday
    Auto-check passed
  • Veomni Review

    ByteDance-Seed/VeOmni

    Pre-PR code review gate. An agent skill from ByteDance-Seed/VeOmni.

    2.2k GitHub stars~1.9k tokensUpdated yesterday
    Auto-check passed

Works with

Questions about Veomni Profile

What does Veomni Profile do?

A skill your agent uses for performance profiling and optimization. Veomni Profile is an agent skill from ByteDance-Seed/VeOmni. Use this skill for performance profiling and optimization.

When should I use Veomni Profile?

Veomni Profile fits situations like: performance profiling and optimization; tasks that involve Performance optimization.

How do I install Veomni Profile in Claude Code?

Run `npx skills add ByteDance-Seed/VeOmni --skill veomni-profile -a claude-code`. Or copy the skill folder (.agents/skills/veomni-profile in ByteDance-Seed/VeOmni) into .claude/skills/veomni-profile in your project. Claude Code loads it when a task matches its description.

How do I install Veomni Profile in Codex?

Run `npx skills add ByteDance-Seed/VeOmni --skill veomni-profile -a codex`. Or copy the skill folder (.agents/skills/veomni-profile in ByteDance-Seed/VeOmni) into .agents/skills/veomni-profile in your project. Codex loads it when a task matches its description.

Can I use Veomni Profile in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ByteDance-Seed/VeOmni --skill veomni-profile -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/veomni-profile, .gemini/skills/veomni-profile, .github/skills/veomni-profile and .opencode/skills/veomni-profile in your project.

What does Veomni Profile need to run?

Going by SKILL.md and its folder, Veomni Profile needs the command-line tools its instructions call (python). Our summary lists: Python 3.

Does Veomni Profile access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Veomni Profile safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Veomni Profile use?

Veomni Profile is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Veomni Profile use?

About 1.7k tokens (SKILL.md is roughly 7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Veomni Profile?

Skills that share tags, products or a category with Veomni Profile: The Art of Debugging (stas00/the-art-of-debugging, 1.7k stars), Cuda (technillogue/ptx-isa-markdown, 229 stars), Torch Profiler Layer Track (BBuf/AI-Infra-Auto-Driven-SKILLS, 938 stars) and Cpp (crazyguitar/cppcheatsheet, 290 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Veomni Profile?

ByteDance-Seed (a GitHub organization) maintains it in ByteDance-Seed/VeOmni, which has 2,235 GitHub stars. The repository holds 10 skills in this directory. The repository was last updated on October 10, 2026.

Source: ByteDance-Seed/VeOmni on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.