Performance Optimization
albumentations-team/AlbumentationsX
Systematic performance audit for AlbumentationsX runtime code.
Route a verl-omni performance investigation to the right tool and capture a usable trace.
$ npx skills add verl-project/verl-omni --skill profile -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install verl-project/verl-omni profile --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/verl-project/verl-omni.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/profile .claude/skills/profile && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "profile" agent skill from https://github.com/verl-project/verl-omni/tree/main/.agents/skills/profile into .claude/skills/profile/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "profile", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/verl-project/verl-omni/tree/main/.agents/skills/profileType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add verl-project/verl-omni --skill profile -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install verl-project/verl-omni profile --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/verl-project/verl-omni.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.agents/skills/profile .agents/skills/profile && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "profile" agent skill from https://github.com/verl-project/verl-omni/tree/main/.agents/skills/profile into .agents/skills/profile/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "profile", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add verl-project/verl-omni --skill profile -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install verl-project/verl-omni profile --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/verl-project/verl-omni.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.agents/skills/profile .cursor/skills/profile && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "profile" agent skill from https://github.com/verl-project/verl-omni/tree/main/.agents/skills/profile into .cursor/skills/profile/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "profile", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/verl-project/verl-omni.git --path .agents/skills/profile--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add verl-project/verl-omni --skill profile -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install verl-project/verl-omni profile --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/verl-project/verl-omni.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.agents/skills/profile .gemini/skills/profile && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "profile" agent skill from https://github.com/verl-project/verl-omni/tree/main/.agents/skills/profile into .gemini/skills/profile/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "profile", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install verl-project/verl-omni profileInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add verl-project/verl-omni --skill profile -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/verl-project/verl-omni.git skills-src && mkdir -p .github/skills && cp -r skills-src/.agents/skills/profile .github/skills/profile && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "profile" agent skill from https://github.com/verl-project/verl-omni/tree/main/.agents/skills/profile into .github/skills/profile/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "profile", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add verl-project/verl-omni --skill profile -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install verl-project/verl-omni profile --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/verl-project/verl-omni.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.agents/skills/profile .opencode/skills/profile && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "profile" agent skill from https://github.com/verl-project/verl-omni/tree/main/.agents/skills/profile into .opencode/skills/profile/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "profile", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
profileRoute a verl-omni performance investigation to the right tool and capture a usable trace.
Profile is an agent skill from verl-project/verl-omni. Route a verl-omni performance investigation to the right tool and capture a usable trace. Use when profiling FlowGRPO / diffusion training or rollout — choosing between nsys, torch.profiler, torchmemory snapshots, MFU comparison, or RL-Insight dashboards, and profiling one lightweight step instead of a full run.
Its SKILL.md is about 1.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in AI & LLM Engineering, covering Performance optimization. The repository describes itself as: Multimodal RL training framework for diffusion & omni models. The licence is Apache-2.0.
3 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 54c557e. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are bash).
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
HF_TOKENFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Profile loads about 1.1k tokens when it runs. Until then it costs about 81 tokens; SKILL.md has 473 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from verl-project/verl-omni at commit 54c557e, republished under its Apache-2.0 licence (© verl-project). 473 words, ~1,088 tokens.
.claude/skills/profile/SKILL.md (or your agent's skills folder).docs/perf/profiler.md owns the config surface (global_profiler +
per-role actor_rollout_ref.{actor,ref,rollout}.profiler), the six copy-paste
recipes, and the lightweight-footprint recipe. Open it — do not work from a
remembered procedure. This skill routes you to the right tool and adds the
cross-cutting decisions the guide leaves implicit.
| The question you are answering | Tool | Where it is documented |
|---|---|---|
| Where does the step's wall-clock go? (phase overlap, Python control flow, rank straggler) | nsys | docs/perf/profiler.md recipes 4, 4a |
| Which ops/kernels dominate, and CPU vs CUDA? | torch (torch.profiler) | docs/perf/profiler.md recipes 1, 2 |
| What is holding GPU memory / who OOMs? | torch_memory | docs/perf/profiler.md recipe 3 |
| Is config B more compute-efficient than A? | MFU (no profiler) | docs/perf/diffusion_mfu.md — read perf/mfu/actor, relative only |
| Live dashboards across replicas / TransferQueue during a long run? | RL-Insight | docs/start/rl_insight.md |
A profiler answers "where is the time/memory in this step". MFU answers "how efficient is this config vs another on the same setup" — it is a metric, not a trace, and it over-estimates LoRA (it counts the full DiT forward+backward).
A full FlowGRPO step trace is hundreds of MB and slow to open. Every examples/
recipe forwards "$@" to the same diffusion_trainer config and Hydra resolves
duplicates last-wins, so append footprint overrides instead of editing the
script — shrink rollout.n, pipeline.num_inference_steps, resolution, and
batch (see the guide's lightweight recipe; it cut a step 616 s → 70 s).
Always pin these, or profiling is silently skipped:
trainer.total_training_steps=1 trainer.save_freq=-1 trainer.test_freq=-1 \
trainer.resume_mode=disable global_profiler.steps=[1]The last step force-triggers save/validation when save_freq/test_freq > 0,
and a leftover checkpoint auto-resumes past the profiled step. For continuous
nsys captures, step 2 is the steady-state sample (step 1 carries profiler
startup, the last step closes the window).
Each phase runs in a different process; enabling the wrong *.profiler yields an
empty trace:
actor_rollout_ref.actor.profileractor_rollout_ref.rollout.profiler (a separate vLLM-Omni server;
tool_config.torch.discrete=True is required — it rejects continuous mode)reward.reward_model.rollout.profileractor_rollout_ref.ref.profilerThese keys already exist in the composed config, so override with plain
key=value — a +key=value append fails with "An item is already at ...".
nsys step-scoped controller capture
(capture-range=cudaProfilerApi) is not supported by
verl_omni.trainer.main_diffusion_v1; only main_diffusion drives the
step-based start/stop lifecycle.nsys output path: *.nsys-rep files land under
/tmp/ray/session_latest/logs/nsight/ (fixed by Ray), not save_path; only
torch / torch_memory traces honor global_profiler.save_path.HF_TOKEN unless
discard-environment is set — scrub before sharing.torch.profiler / timers inside adapter or pipeline code.
Workers are wrapped with verl.utils.profiler.DistProfiler and driven around
each profiled step; an ad-hoc timer is fine for a throwaway local check but
must never reach a PR.docs/perf/profiler.md — authoritative config surface and recipes.docs/perf/diffusion_mfu.md — MFU reporting and adding an estimator.docs/start/rl_insight.md — online observability dashboards.© verl-project, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .agents/skills/profile of verl-project/verl-omni.
Open the folder on GitHubat commit 54c557e
Profile next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Profile this skillverl-project/verl-omni | 1.2k | — | ~1.1k | Automated safety check: Pass | Apache-2.0 | |
| Performance Optimizationalbumentations-team/AlbumentationsX | 567 | — | ~1.7k | Automated safety check: Pass | AGPL-3.0 | |
| Optimizing Promptsjeremylongshore/tons-of-skills-marketplace | 2.8k | — | ~1k | Automated safety check: Pass | MIT | |
| Perfupraullenchai/Rapid-MLX | 3.9k | — | ~1.6k | Automated safety check: Notes | Custom licence | |
| The Art of Debuggingstas00/the-art-of-debugging | 1.7k | — | ~6.1k | Automated safety check: Notes | CC-BY-SA-4.0 | |
| LLM Torch Profiler Trace AnalysisBBuf/AI-Infra-Auto-Driven-SKILLS | 925 | — | ~2.8k | Automated safety check: Pass | None |
albumentations-team/AlbumentationsX
Systematic performance audit for AlbumentationsX runtime code.
jeremylongshore/tons-of-skills-marketplace
Execute this skill optimizes prompts for large language models (llms) to reduce token usage, lower costs, and improve performance.
raullenchai/Rapid-MLX
Autonomous performance optimization: research, PoC, benchmark, implement, review, PR
stas00/the-art-of-debugging
Condensed debugging method and tool recipes for Unix, Python and PyTorch programs: crashes, hangs, segfaults, wrong output, CUDA OOM, NaN values and slowness.
BBuf/AI-Infra-Auto-Driven-SKILLS
Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.
sgl-project/sglang
A skill your agent uses when benchmarking denoise latency or profiling a diffusion bottleneck in SGLang.
verl-project/verl-omni
Router for adding a diffusion or omni pipeline to verl-omni.
verl-project/verl-omni
Guide for adding a new reward scorer to verl-omni and wiring it into a run.
verl-project/verl-omni
How to write and run verl-omni CPU tests (testoncpu.py) that exercise adapters, rewards, and configs without a GPU or model weights.
verl-project/verl-omni
verl-omni commit message + PR conventions and the mandatory contribution policy.
verl-project/verl-omni
Review your own verl-omni branch against the project rubric before opening or updating a PR.
verl-project/verl-omni
Route verl-omni training/inference consistency checks through MindStudio's MSProbe collection and root-cause analysis skills.
Categories
Route a verl-omni performance investigation to the right tool and capture a usable trace. Profile is an agent skill from verl-project/verl-omni. Route a verl-omni performance investigation to the right tool and capture a usable trace.
Profile fits situations like: profiling FlowGRPO / diffusion training; rollout — choosing between nsys; torchmemory snapshots; RL-Insight dashboards.
Run `npx skills add verl-project/verl-omni --skill profile -a claude-code`. Or copy the skill folder (.agents/skills/profile in verl-project/verl-omni) into .claude/skills/profile in your project. Claude Code loads it when a task matches its description.
Run `npx skills add verl-project/verl-omni --skill profile -a codex`. Or copy the skill folder (.agents/skills/profile in verl-project/verl-omni) into .agents/skills/profile in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add verl-project/verl-omni --skill profile -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/profile, .gemini/skills/profile, .github/skills/profile and .opencode/skills/profile in your project.
Going by SKILL.md and its folder, Profile needs credentials named HF_TOKEN. Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Profile is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.1k tokens (SKILL.md is roughly 4.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Profile: Performance Optimization (albumentations-team/AlbumentationsX, 567 stars), Optimizing Prompts (jeremylongshore/tons-of-skills-marketplace, 2.8k stars), Perfup (raullenchai/Rapid-MLX, 3.9k stars) and The Art of Debugging (stas00/the-art-of-debugging, 1.7k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
verl-project (a GitHub organization) maintains it in verl-project/verl-omni, which has 1,204 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on October 9, 2026.
Source: verl-project/verl-omni on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.