Astrea
warpfront/hipfire
A skill your agent uses for hipfire quant calibration, imatrix-driven experiments, KLD/PPL quality evaluation, k-map/format selection, MQ/HFQ/HFP/MFP tradeoff work, ParoQuant-style weight transform…
Analyze and optimize vllm-rlt inference performance using reproducible unprofiled benchmarks, paired ops-only/full profiles, source-level attribution, and correctness checks.
$ npx skills add ThinkFlowLab/vllm-rlt --skill rlt-perf-opt -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install ThinkFlowLab/vllm-rlt rlt-perf-opt --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/ThinkFlowLab/vllm-rlt.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/rlt-perf-opt .claude/skills/rlt-perf-opt && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "rlt-perf-opt" agent skill from https://github.com/ThinkFlowLab/vllm-rlt/tree/main/.agents/skills/rlt-perf-opt into .claude/skills/rlt-perf-opt/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "rlt-perf-opt", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/ThinkFlowLab/vllm-rlt/tree/main/.agents/skills/rlt-perf-optType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add ThinkFlowLab/vllm-rlt --skill rlt-perf-opt -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install ThinkFlowLab/vllm-rlt rlt-perf-opt --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ThinkFlowLab/vllm-rlt.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.agents/skills/rlt-perf-opt .agents/skills/rlt-perf-opt && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "rlt-perf-opt" agent skill from https://github.com/ThinkFlowLab/vllm-rlt/tree/main/.agents/skills/rlt-perf-opt into .agents/skills/rlt-perf-opt/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "rlt-perf-opt", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ThinkFlowLab/vllm-rlt --skill rlt-perf-opt -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install ThinkFlowLab/vllm-rlt rlt-perf-opt --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ThinkFlowLab/vllm-rlt.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.agents/skills/rlt-perf-opt .cursor/skills/rlt-perf-opt && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "rlt-perf-opt" agent skill from https://github.com/ThinkFlowLab/vllm-rlt/tree/main/.agents/skills/rlt-perf-opt into .cursor/skills/rlt-perf-opt/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "rlt-perf-opt", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/ThinkFlowLab/vllm-rlt.git --path .agents/skills/rlt-perf-opt--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add ThinkFlowLab/vllm-rlt --skill rlt-perf-opt -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install ThinkFlowLab/vllm-rlt rlt-perf-opt --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ThinkFlowLab/vllm-rlt.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.agents/skills/rlt-perf-opt .gemini/skills/rlt-perf-opt && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "rlt-perf-opt" agent skill from https://github.com/ThinkFlowLab/vllm-rlt/tree/main/.agents/skills/rlt-perf-opt into .gemini/skills/rlt-perf-opt/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "rlt-perf-opt", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install ThinkFlowLab/vllm-rlt rlt-perf-optInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add ThinkFlowLab/vllm-rlt --skill rlt-perf-opt -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/ThinkFlowLab/vllm-rlt.git skills-src && mkdir -p .github/skills && cp -r skills-src/.agents/skills/rlt-perf-opt .github/skills/rlt-perf-opt && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "rlt-perf-opt" agent skill from https://github.com/ThinkFlowLab/vllm-rlt/tree/main/.agents/skills/rlt-perf-opt into .github/skills/rlt-perf-opt/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "rlt-perf-opt", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ThinkFlowLab/vllm-rlt --skill rlt-perf-opt -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install ThinkFlowLab/vllm-rlt rlt-perf-opt --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ThinkFlowLab/vllm-rlt.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.agents/skills/rlt-perf-opt .opencode/skills/rlt-perf-opt && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "rlt-perf-opt" agent skill from https://github.com/ThinkFlowLab/vllm-rlt/tree/main/.agents/skills/rlt-perf-opt into .opencode/skills/rlt-perf-opt/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "rlt-perf-opt", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
rlt-perf-optAnalyze and optimize vllm-rlt inference performance using reproducible unprofiled benchmarks, paired ops-only/full profiles, source-level attribution, and correctness checks.
Rlt Perf Opt is an agent skill from ThinkFlowLab/vllm-rlt. Analyze and optimize vllm-rlt inference performance using reproducible unprofiled benchmarks, paired ops-only/full profiles, source-level attribution, and correctness checks. Use for quantitative PR performance reviews and optimization of looped Transformer inference features.
Its SKILL.md is about 4.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in AI & LLM Engineering, covering LLM inference and serving and Performance reviews. It works with vLLM. The repository describes itself as: vllm based inference engine for looped transformer. The licence is Apache-2.0.
8 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit b599dc5. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md.
From the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
github.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Rlt Perf Opt loads about 4.6k tokens when it runs. Until then it costs about 73 tokens; SKILL.md has 2,427 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from ThinkFlowLab/vllm-rlt at commit b599dc5, republished under its Apache-2.0 licence (© ThinkFlowLab). 2,427 words, ~4,630 tokens.
.claude/skills/rlt-perf-opt/SKILL.md (or your agent's skills folder).The goal is to answer: how does the feature work, which code or execution stages account for performance changes, how do those changes arise quantitatively, and which further optimizations are worth implementing? Final conclusions must be verifiable from configurations, raw results, traces, and source locations.
This skill covers performance analysis and optimization after feature implementation. Code-style refactoring belongs to a separate workflow; do not mix unrelated refactoring into performance experiments. When the user requests analysis only, deliver measurements and recommendations; implementation changes must stay within the user's authorization.
Before measuring, use the requirements, PR diff, and actual code to explain:
A line-by-line walkthrough is unnecessary, but do not infer behavior from a feature name alone. For unsupported combinations, record the current code version, the source location checked or minimal reproduction results, and mark the combination as untested. Do not combine gains from two independent tests and present them as a result for the combined mode.
First read the project's existing benchmarks, user guides, and profiling interfaces, and prefer reusing existing tools. Put any missing diagnostic scripts in a local experiment directory; do not add a new benchmark framework to the repository by default.
Save three types of entry points: service/engine startup commands, single-request reproduction commands, and batch benchmark commands. Associate each experiment with the following information:
State whether the experiment isolates the cost of a particular feature or compares the best tuned configurations of different modes. Do not mix these two types of results in the same performance-gain claim. Explain any configuration differences that cannot be held constant.
Resource fairness: compare aggregate throughput for a PD 1P1D deployment against a conventional inference deployment using the same two GPUs. For single requests, an additional comparison against single-GPU latency is acceptable, but state the resource cost. For speculative versus target-only inference, hold the model, target exit policy, and hardware budget fixed.
Choose based on the mechanism being tested; no particular dataset is mandatory:
Optional starting scenarios: 128 input tokens / 64 output tokens / one request; 1024 input tokens / 64 output tokens / 16 simultaneous arrivals; 4096 input tokens / 64 output tokens / fixed-rate arrivals. Adjust to the objective. Short bursts cannot establish steady-state throughput or saturation capacity. High-concurrency conclusions require a load sweep and a sufficiently long stable measurement window.
Record fixed-shape ignore-EOS experiments separately from natural generation. Also separate cold/warm cache conditions, graph compilation/capture, and steady-state execution.
Performance numbers must come from independent runs with the profiler disabled. Warmup must cover the actual shape, batch, and graph paths. By default, collect at least three valid repetitions and retain each run's results and variability. For noise-sensitive or small gains, check for drift using methods such as interleaved A/B runs.
Record at least:
| Metric | Definition and reporting requirements |
|---|---|
| Output throughput | Successfully generated output tokens / an explicitly defined measurement window; also report total output tokens, window duration, and GPU count |
| Request throughput | Successful requests / the same window |
| TTFT | Time from the client sending the request to the first valid output; distinguish server-side queuing and network timing boundaries |
| TPOT | Specify the first-to-last token time difference and denominator used; explain observation limitations for speculative output emitted in batches |
| ITL / chunk gap | Use ITL only when per-token timestamps are available; SSE chunk intervals must not be presented as speculative per-token intervals |
| E2E latency | Time from sending a single request to its completion |
| Distributions and failures | Latency p50/p95/p99, sample counts, and error/timeout/cancellation rates; tail percentiles from small samples are descriptive only |
| Resource usage | Relevant peak device memory, GPU/CPU activity, KV usage, and transfer volume; state the measurement method |
When SLOs apply, also report goodput subject to TTFT/TPOT constraints. Do not silently exclude failed requests. Means, medians of per-run percentiles, and percentiles of pooled samples are different statistics; specify which is being reported.
Independently capture both modes using the same representative request configuration, with fresh sessions and output directories:
| Setting | ops-only | full |
|---|---|---|
| CPU + device operator activity | On | On |
| record_shapes | Off | On |
| profile_memory | Off | On |
| with_stack | Off | On |
| with_flops | Off by default | Off by default; enable only when needed |
Use ops-only to identify operators, launches, copies, synchronization, and device timelines. Use full to further locate shapes, allocations, and call stacks. Full does not mean enabling only shapes or only stacks. On non-CUDA platforms, use the device profiler actually supported by that platform and disclose missing capabilities; do not describe a CPU-only capture as a complete device capture.
Start with the smallest window that covers the mechanism under investigation, usually one complete representative speculative round or decode iteration after warmup. An engine step is not necessarily a complete token iteration; ensure the window includes the relevant draft, verification, commit/rollback, or transfer stages. Increase the window only when the initial trace lacks necessary evidence. Benchmark repetition counts do not determine profiling step counts.
Inspect the first capture's file size, event count, export time, and parsing cost before launching a profiling matrix. If traces become too large, export or analysis takes too long, or capture/export fails, first reduce the number of active steps or narrow the capture to the relevant stage. Keep the request configuration representative, and use matching windows for the baseline and candidate and for ops-only/full. Split distinct stages into separate short captures when needed. Do not repeatedly rerun the same oversized window or broaden the matrix before validating a small capture.
Reduce the window, not the definition of full: keep record_shapes, profile_memory, and with_stack enabled. Preserve enough events to support the conclusion; an incomplete round cannot stand in for the whole round. If reducing the window does not resolve an export error, retain the error and then investigate profiler/runtime compatibility. Do not assume file size caused an error without evidence.
Inspect summaries and selected events before loading an entire large trace. Preserve existing valid captures and reuse them; do not recapture completed evidence merely to obtain smaller files. Report unprofiled A/B results as soon as they are available, with their validation status, rather than withholding all results while profiling or export issues are being resolved.
Retain traces, operator statistics, metadata, capture commands, and corresponding requests. Do not commit large traces directly to the code repository.
Start from the end-to-end difference, then examine: request queuing and batch composition → prefill/loops/decode → host scheduling and device idle gaps → operator hotspots, memory, and transfers. Link each finding to file:line, the call path, trace events, and the measurement definition.
Distinguish carefully between:
.cpu(), .item(), or synchronization APIs and actual D2H transfer time. A wait may include preceding GPU computation; do not treat the entire wait as an eliminable copy.Provide an evidence chain for each major gain or regression:
configuration change → code/flow change → change in work, event counts, or waits → unprofiled metric difference → added costs and remaining uncertainty
Where quantification is possible, report both absolute differences and relative changes. If the current data cannot isolate a causal contribution, explicitly label it as a hypothesis and design an A/B test, ablation, or additional counter to validate it. Do not force contributions into an explanation that adds up to 100%.
Prioritize candidates by critical-path cost, expected gain, complexity, and correctness risk. Distinguish proven bottlenecks, hypotheses awaiting validation, and measured optimizations.
Where possible, change one interpretable factor per iteration and record the code/configuration differences before and after optimization. First examine existing capabilities: backends, batch/token budgets, graph coverage, redundant synchronization/copies, repeated computation, and metadata preparation. Introduce new kernels or complex scheduling only when supported by evidence.
Once implementation is authorized, iterate: implement → check correctness → run unprofiled A/B measurements under the same definitions → reprofile when necessary → decide whether to retain, revert, or continue. Analysis-only tasks may deliver concrete change proposals and validation experiments; do not expand them into implementation without authorization.
Do not change request semantics, exit policies, or accuracy for performance without disclosure. Report gains from shorter outputs, fewer loops, or different sampling algorithms separately as quality/performance tradeoffs.
Deliver correctness results alongside performance results:
If unexplained correctness differences remain, performance numbers may be retained as diagnostic results, but the optimization must not be declared to have passed acceptance.
The report must let readers understand the conclusions without rerunning the experiments and reproduce them when needed. Include:
Describe the best result within the current hardware, model, and tested cases only as “best among tested configurations” or “the most suitable tradeoff under current conditions.” State which dimensions were not searched; do not claim a global optimum.
After correctness checks pass and performance conclusions stabilize, collect the configurations, commands, metrics, versions, and applicability limits needed for a recipe. If a recipe already exists, update it using the repository's format. Otherwise, first save reusable materials locally; publication must follow the current task's authorization. New features, parameters, or usage patterns also require updates to the relevant documentation and README.
This skill draws on the reproducible baselines, profiling-based attribution, and iterative validation approach of vllm-omni diffusion-perf-opt, reorganized for looped Transformers. It does not incorporate diffusion-specific parallel-strategy search. The definitions of ops-only and full in this skill take precedence.
© ThinkFlowLab, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .agents/skills/rlt-perf-opt of ThinkFlowLab/vllm-rlt.
Open the folder on GitHubat commit b599dc5
Rlt Perf Opt next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Rlt Perf Opt this skillThinkFlowLab/vllm-rlt | 149 | — | ~4.6k | Automated safety check: Pass | Apache-2.0 | |
| Astreawarpfront/hipfire | 658 | — | ~2.6k | Automated safety check: Pass | Custom licence | |
| Quark Onnx Autosearch Proamd/Quark | 182 | — | ~3.4k | Automated safety check: Pass | MIT | |
| Quark Onnx Quant Planamd/Quark | 182 | — | ~4.8k | Automated safety check: Pass | MIT | |
| SageMaker Serving Image Selectionhuggingface/skills | 11k | 1 repos | ~4.6k | Automated safety check: Pass | Apache-2.0 | |
| Hugging Face Local Model Evalshuggingface/skills | 11k | 2 repos | ~1.6k | Automated safety check: Pass | Apache-2.0 |
warpfront/hipfire
A skill your agent uses for hipfire quant calibration, imatrix-driven experiments, KLD/PPL quality evaluation, k-map/format selection, MQ/HFQ/HFP/MFP tradeoff work, ParoQuant-style weight transform…
amd/Quark
L3 recipe that runs quark.onnx.AutoSearchPro end-to-end on a user .onnx model: intake → preset selection (or custom search space) → calibration / eval data reader → standalone autosearch script…
amd/Quark
Build a Quark ONNX PTQ quantization plan from modelanalysis.json and user intent.
huggingface/skills
Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.
huggingface/skills
Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.
guqiong96/Lvllm
Fetch and diagnose vLLM Buildkite CI failure logs. An agent skill from guqiong96/Lvllm.
ThinkFlowLab/vllm-rlt
Review and refactor inference-runtime code using concrete rules for responsibility boundaries, state ownership, interfaces, asynchronous lifetimes, KV management, and maintainability.
ThinkFlowLab/vllm-rlt
Review PRs and local changes for hsliuustc0106/vllm-rlt: Ouro engine and KV correctness, serving behavior, and BF16 accuracy/speed evidence.
Works with
Analyze and optimize vllm-rlt inference performance using reproducible unprofiled benchmarks, paired ops-only/full profiles, source-level attribution, and correctness checks. Rlt Perf Opt is an agent skill from ThinkFlowLab/vllm-rlt. Analyze and optimize vllm-rlt inference performance using reproducible unprofiled benchmarks, paired ops-only/full profiles, source-level attribution, and correctness checks.
Rlt Perf Opt fits situations like: quantitative PR performance reviews and optimization of looped Transformer inference features; tasks that involve LLM inference and serving; tasks that involve Performance reviews.
Run `npx skills add ThinkFlowLab/vllm-rlt --skill rlt-perf-opt -a claude-code`. Or copy the skill folder (.agents/skills/rlt-perf-opt in ThinkFlowLab/vllm-rlt) into .claude/skills/rlt-perf-opt in your project. Claude Code loads it when a task matches its description.
Run `npx skills add ThinkFlowLab/vllm-rlt --skill rlt-perf-opt -a codex`. Or copy the skill folder (.agents/skills/rlt-perf-opt in ThinkFlowLab/vllm-rlt) into .agents/skills/rlt-perf-opt in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ThinkFlowLab/vllm-rlt --skill rlt-perf-opt -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/rlt-perf-opt, .gemini/skills/rlt-perf-opt, .github/skills/rlt-perf-opt and .opencode/skills/rlt-perf-opt in your project.
SKILL.md names no scripts, command-line tools or credentials: Rlt Perf Opt is instructions for the agent only.
SKILL.md names 1 domain. As links in the text: github.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Rlt Perf Opt is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.6k tokens (SKILL.md is roughly 19k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Rlt Perf Opt: Astrea (warpfront/hipfire, 658 stars), Quark Onnx Autosearch Pro (amd/Quark, 182 stars), Quark Onnx Quant Plan (amd/Quark, 182 stars) and SageMaker Serving Image Selection (huggingface/skills, 11k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
ThinkFlowLab (a GitHub organization) maintains it in ThinkFlowLab/vllm-rlt, which has 149 GitHub stars. The repository holds 3 skills in this directory. The repository was last updated on October 11, 2026.
Source: ThinkFlowLab/vllm-rlt on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.