Matlab Use Visual Inspection
matlab/matlab-agentic-toolkit
Build machine vision inspection systems with MATLAB Visual Inspection Toolbox.
Explains where HOT-Step generation time goes (LM/DiT/VAE), how the TensorRT paths activate, how to benchmark from logs, and which knobs trade quality for speed.
$ npx skills add scragnog/HOT-Step-CPP --skill engine-performance -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install scragnog/HOT-Step-CPP engine-performance --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/scragnog/HOT-Step-CPP.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/engine-performance .claude/skills/engine-performance && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "engine-performance" agent skill from https://github.com/scragnog/HOT-Step-CPP/tree/master/.claude/skills/engine-performance into .claude/skills/engine-performance/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "engine-performance", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/scragnog/HOT-Step-CPP/tree/master/.claude/skills/engine-performanceType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add scragnog/HOT-Step-CPP --skill engine-performance -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install scragnog/HOT-Step-CPP engine-performance --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/scragnog/HOT-Step-CPP.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills/engine-performance .agents/skills/engine-performance && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "engine-performance" agent skill from https://github.com/scragnog/HOT-Step-CPP/tree/master/.claude/skills/engine-performance into .agents/skills/engine-performance/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "engine-performance", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add scragnog/HOT-Step-CPP --skill engine-performance -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install scragnog/HOT-Step-CPP engine-performance --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/scragnog/HOT-Step-CPP.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills/engine-performance .cursor/skills/engine-performance && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "engine-performance" agent skill from https://github.com/scragnog/HOT-Step-CPP/tree/master/.claude/skills/engine-performance into .cursor/skills/engine-performance/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "engine-performance", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/scragnog/HOT-Step-CPP.git --path .claude/skills/engine-performance--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add scragnog/HOT-Step-CPP --skill engine-performance -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install scragnog/HOT-Step-CPP engine-performance --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/scragnog/HOT-Step-CPP.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills/engine-performance .gemini/skills/engine-performance && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "engine-performance" agent skill from https://github.com/scragnog/HOT-Step-CPP/tree/master/.claude/skills/engine-performance into .gemini/skills/engine-performance/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "engine-performance", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install scragnog/HOT-Step-CPP engine-performanceInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add scragnog/HOT-Step-CPP --skill engine-performance -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/scragnog/HOT-Step-CPP.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills/engine-performance .github/skills/engine-performance && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "engine-performance" agent skill from https://github.com/scragnog/HOT-Step-CPP/tree/master/.claude/skills/engine-performance into .github/skills/engine-performance/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "engine-performance", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add scragnog/HOT-Step-CPP --skill engine-performance -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install scragnog/HOT-Step-CPP engine-performance --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/scragnog/HOT-Step-CPP.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills/engine-performance .opencode/skills/engine-performance && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "engine-performance" agent skill from https://github.com/scragnog/HOT-Step-CPP/tree/master/.claude/skills/engine-performance into .opencode/skills/engine-performance/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "engine-performance", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
engine-performanceExplains where HOT-Step generation time goes (LM/DiT/VAE), how the TensorRT paths activate, how to benchmark from logs, and which knobs trade quality for speed.
Engine Performance is an agent skill from scragnog/HOT-Step-CPP. Explains where HOT-Step generation time goes (LM/DiT/VAE), how the TensorRT paths activate, how to benchmark from logs, and which knobs trade quality for speed. Use when profiling slow generations, working on dit-trt.h/lm-trt.h/adapter-trt.h/hot-step-sampler-trt.h, debugging TRT/ONNX issues, or deciding what performance work is DONE vs PLANNED.
Its SKILL.md is about 4.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `reference.md`).
It sits in AI & LLM Engineering, covering LLM inference and serving. It works with NVIDIA AI Platform, ONNX and C++. The repository describes itself as: Turn dials. Summon bangers! NOW WITH MORE C++! Local AI music generation powered by GGML. The licence is MIT.
8 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 91e92a8. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
cmakeFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Engine Performance loads about 4.9k tokens when it runs. Until then it costs about 91 tokens; SKILL.md has 2,329 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from scragnog/HOT-Step-CPP at commit 91e92a8, republished under its MIT licence (© scragnog). 2,329 words, ~4,925 tokens.
.claude/skills/engine-performance/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.The C++ engine (engine/) generates music in three phases: LM (a Qwen3 language model produces audio-code tokens from caption+lyrics), DiT (a Diffusion Transformer denoises latents over N steps — this is the main compute loop), and VAE (decodes latents to 48 kHz stereo audio). Two inference backends coexist: GGML (quantized GGUF files, the default) and TensorRT ("TRT" — NVIDIA's compiled-engine runtime, fed by ONNX model exports). This skill maps the performance landscape, the TRT integration, and how to measure things.
Detailed line-references, the full DONE/PLANNED ledger, and the on-disk model layout live in reference.md.
engine/src/dit-trt.h, lm-trt.h, adapter-trt.h, hot-step-sampler-trt.h, lm-trtllm.h, stream-pipeline.h, or the ORT wrappers (vae-ort.h, text-enc-ort.h, cond-enc-ort.h, vae-enc-ort.h).dev-rebuild.bat at repo root, NEVER engine/build.cmd directly. WHY: you cannot reliably tell whether the app is running; the Node server auto-respawns ace-server.exe on crash, and killing it mid-build causes an infinite respawn + file-lock loop.cmake --build . --clean-first unless the GGML/CUDA layer itself changed. WHY: CUDA kernel recompilation takes 20+ minutes. For stale .obj issues delete only engine/build/acestep-core.dir/ and engine/build/Release/acestep-core.lib..engine files between machines. WHY: TRT engines are locked to the GPU architecture (sm_86/89/120 are not interchangeable) and TRT version. They are built on first use and cached next to the ONNX file — that is the intended distribution model..engine file's sibling .onnx. WHY: engines are built with kSTRIP_PLAN (weights stripped, ~50 MB plan); weights are re-refit from the ONNX on every load (engine/src/dit-trt.h:375-377). Engine without ONNX = unusable.lm_trt_free (engine/src/lm-trt.h:702-709: cudaStreamSynchronize + cudaDeviceSynchronize before freeing). WHY: TRT teardown without them corrupts shared CUDA state and crashes GGML afterwards — this was learned from real crashes.[Adapter-TRT] Applied in N ms, [DiT-Generate] TRT Total: ....engine/src/hot-step-sampler.h:593 — "Confirmed: skipping these produces blank output." The GGML scheduler aliases input buffers as scratch, so encoder/pos/mask constants must be re-uploaded every step. The TRT sampler re-uploads enc_hidden per step for the same class of reason. A local design doc marks this optimization "implemented"; the code comment wins.Measured breakdown from the local TRT optimization plan (RTX 5090, 4B LM + XL DiT, 50 steps, same song/adapter — predates the adapter-batching fix and dedicated CUDA streams, so treat as directional):
| Component | TRT | GGUF | Winner |
|---|---|---|---|
| LM Phase | 29.4s | 18.0s | GGUF |
| Adapter apply | 16.6s → later fixed to ~1.8s | 2.0s | ~tie now |
| DiT Denoising | 15.3s | 20.5s | TRT (~−25%) |
| DiT Model Load | 7.5s | ~0s | GGUF |
| VAE Decode | 0.8s | 1.3s | TRT (~−38%) |
| Text Encoding | ~6.4s | ~5s | ~tie |
Rules of thumb: DiT denoising dominates (turbo models = 8 steps, base/SFT = 50 steps, cost is linear in steps). Per-step GGML DiT cost is ~85–90% GPU transformer layers, ~5–8% sync/re-upload overhead, ~2–3% CPU-side guidance, <1% solver. TRT wins the DiT and VAE; GGUF wins the LM and load times. No fresher end-to-end A/B exists in the repo (unverified whether TRT now wins overall).
Two distinct TRT integrations coexist:
IRefitter, which runtime LoRA/LoKr adapter switching requires (engine/src/dit-trt.h header comment).engine/src/vae-ort.h:78).Triggers:
.onnx → dit_ends_with_onnx gate at engine/src/pipeline-synth-ops.cpp:1225 → dit-trt.h. Engine file = same name with .engine; built on first use if missing (5–30 min), else loaded from cache.<model_path>/lm_full.onnx exists → raw TRT LM (lm-trt.h), engine lm_full.engine built if missing (engine/src/pipeline-lm.cpp:1159). A TRT-LLM engine dir (trtllm-engine-*) is checked first but TRT-LLM is compile-time disabled (see ledger).use_ort_vae request flag are gone; the registry takes the first non-ONNX VAE (registry_find_non_onnx, engine/tools/hot-step-server.cpp:1311). There is no useOrtVae in translateParams.ts any more.models/onnx/pp-vae/ — no user action.stream_mode request): requires an ONNX DiT (pipeline-synth-ops.cpp:2098); ring-buffer batched generation, previews currently hard-disabled (:2251 sets preview_interval = 0 — partially-denoised latents through the VAE = scrambled audio).Build requirement: the vendored TRT SDK at engine/deps/tensorrt/ (present: include/NvInfer.h, libs *_10 = TRT 10.x) defines HOT_STEP_TRT at compile time (engine/CMakeLists.txt:302). Without it, all TRT code compiles out and the build is GGML-only. ONNX export tooling: tools/onnx-export/ (export_dit.py, export_lm.py, export_vae.py, etc.) — runs in a separate Python venv (hot-step-9000 repo), not in this repo's toolchain.
There is no dedicated bench tool — instrumentation is log-based. Logs land in logs\YYYY-MM-DD_HH-MM-SS\ per session (name-sorted = time-sorted).
http://localhost:3000 (app started detached with dev.bat).logs\<newest>\generations\gen_<uuid>_<task>.log as a [Timing] ── Pipeline Breakdown ── block (rendered by server/src/routes/generate.ts:1257) with per-stage seconds/percent bars: LM Phase, model loads, FSQ Detokenize, Adapter Refit/Merge, DiT Denoising, VAE Decode, Text Encoding, plus gap rows (HTTP→Engine, DiT Model Load, etc.).logs\<newest>\ace_engine.log (see Golden rule 6). Key markers:[DiT-Generate] TRT Total: X ms (Y ms/sample) (pipeline-synth-ops.cpp:1384) or the GGML equivalent [DiT-Generate] Total: ...[DiT-TRT] Step k/N ... (first step: N ms) — first-step latency including warmup (hot-step-sampler-trt.h:700)[Adapter-TRT] Applied in N ms / Batched GPU merge: N ms[DiT-TRT] Load + refit complete (N ms); [LM-TRT] Load complete ...[VAE-Decode ...] Decode: N ms (ORT|GGML)$s = Get-ChildItem D:\Ace-Step-Latest\hot-step-cpp\logs | Sort-Object Name -Descending | Select-Object -First 1
Select-String -Path "$($s.FullName)\ace_engine.log" -Pattern 'TRT Total|Applied in|Decode:|Load complete|first step'
Select-String -Path "$($s.FullName)\generations\*.log" -Pattern '\[Timing\]'.engine cache exists.| Knob | Where | Trade-off |
|---|---|---|
CFG Cutoff (cfg_cutoff_ratio, default 1.0) | UI Performance accordion (GenerationDropdown.tsx); engine hot-step-params.h:170 | Classifier-free guidance (the 2× "conditional + unconditional" forward pass) runs only for the first ratio×steps, then conditional-only — halves per-step cost after the cutoff. ~0.5 ≈ 20% speedup; may reduce prompt adherence. TRT path also frees/halves GPU buffers at cutoff. |
LM CFG Cutoff (lm_cfg_cutoff_ratio) | request.h, pipeline-lm.cpp | Same idea for LM token generation; ~0.7 ≈ 15% LM speedup. |
Step Cache (cache_ratio, default 0.0, UI max 0.7) | hot-step-params.h:176; GGML hot-step-sampler.h; TRT hot-step-sampler-trt.h:484 | Skips middle-step forward passes, reusing last velocity; first/last steps protected. Try 0.3–0.5. Stacks with CFG cutoff. |
| Steps / model tier | num_steps; turbo (8) vs base/SFT (50) | Linear cost. Biggest single lever. |
| Quantization | GGUF variants in models\ (Q4_K_M → Q8_0 → BF16, plus NVFP4/MXFP4) | VRAM vs dequant overhead vs quality. |
Co-resident models (coResident → engine EVICT_NEVER) | generate.ts:302/584; pipeline-synth-ops.cpp:1390 | Keeps DiT (including multi-GB TRT engine) + VAE in VRAM between jobs — saves ~7.5s reload per back-to-back run, costs GBs of VRAM. Default EVICT_STRICT frees after every generation. |
Batched CFG (use_batch_cfg) | request | One 2N-batch forward vs two N passes; faster but doubles activation VRAM. |
Flash attention (use_fa) | request | GGML paths. |
VAE tiling (vae_chunk/vae_overlap, defaults 256/64) | synth params | Bigger tiles = fewer passes and fewer seams, more VRAM. |
LM draft model (--draft-model) | pipeline-lm.cpp:793 speculative decode | 0.6B draft proposes, 4B verifies. GGML-side LM speedup. |
Streaming (stream_mode/stream_depth) | request; TRT-only | Ring-buffer batched denoising; previews disabled pending temporal chunking. |
EVICT_NEVER.LM_TRT_VOCAB 217204, lm-trt.h:42); prefill optimization profile; explicit TRT warmup call; GGML CUDA graph capture (no GGML_CUDA_GRAPH anywhere in engine source); SageAttention; real streaming previews; adapter support on FP8/fp32-I/O engines.engine/CMakeLists.txt:344, "Native Windows TRT-LLM is not viable"; TRT version mismatch between Docker-built engines and Windows SDK). Code kept behind -DHOT_STEP_TRTLLM_ENABLE=ON for a future WSL2 path.| Path | Role |
|---|---|
engine/src/dit-trt.h | Native TRT DiT: engine build (dit_trt_build:198), load+refit (dit_trt_load:377), adapter refit (:577), forward (:674) |
engine/src/hot-step-sampler-trt.h | TRT denoising loop — mirrors the GGML sampler (hot-step-sampler.h), full plugin parity; batched streaming step (dit_trt_step:739) |
engine/src/adapter-trt.h | LoRA/LoKr → TRT refit; batched GPU LoKr merge (entry adapter_trt_apply:604) |
engine/src/lm-trt.h | Raw TRT Qwen3-4B LM, double-buffered KV cache, teardown syncs at :702-709 |
engine/src/lm-trtllm.h | TRT-LLM Executor — compile-time disabled, do not assume it runs |
engine/src/stream-pipeline.h | Ring-buffer streaming generation (TRT-only) |
engine/src/vae-ort.h, text-enc-ort.h, cond-enc-ort.h, vae-enc-ort.h | ORT + TRT-EP wrappers for refit-free models |
engine/src/pipeline-synth-ops.cpp | Backend dispatch: TRT-vs-GGML gate :1225, eviction policy :1390, ORT VAE gate :1589, streaming :2095 |
engine/src/pipeline-lm.cpp | LM backend dispatch (TRT-LLM → raw TRT → GGML probe order :1110-1193), speculative decoding :790 |
engine/src/hot-step-params.h | cfg_cutoff_ratio:170, cache_ratio:176 |
engine/tools/hot-step-server.cpp | CLI parsing; registry scan (including the onnx/ subdirectory) and PP-VAE auto-detect |
engine/CMakeLists.txt | HOT_STEP_TRT define :302; TRT-LLM disable block :342-371 |
server/src/routes/generate.ts | Timing table :1257, stream markers :787-798, coResident :302 |
server/src/services/generation/translateParams.ts | UI param → engine request mapping (useOrtVae:186) |
tools/onnx-export/ | Python ONNX exporters (separate venv, not this repo's toolchain) |
docs/plans/2026-05-31-TRT-OPTIMIZATION-PLAN.md | Local-only optimization plan + measurements (gitignored — may be absent) |
| Symptom | Cause → fix |
|---|---|
First generation with an ONNX model "hangs" 5–30 min; log shows [DiT-TRT] This will take 5-30 minutes + Engine build in progress... (Ns elapsed) every 30s | Normal first-run engine build; result cached as .engine next to the ONNX. Deleting .engine retriggers it. The heartbeat exists so the Node stall-detector doesn't kill the job. |
| TRT loads but garbage/NaN audio | Check [DiT-TRT] DIAG first output: (dit-trt.h:743) for NaNs — ONNX export precision issue (norms must be fp32; kSTRONGLY_TYPED forbids per-layer precision overrides, so a re-export is required). |
| Adapter has no audible effect on TRT | Log shows either [Adapter-TRT] No weights matched TRT engine (name-mapping failure) or the FP8/fp32-engine "adapters not supported with this I/O dtype" warning (pipeline-synth-ops.cpp near :1320). |
| Adapter deltas subtly wrong on TRT only | Missing <onnx>.refit_manifest.json sidecar — dynamo-transposed weights get deltas in the wrong orientation. Log: No refit manifest found. |
| GGML CUDA crash right after LM-TRT teardown | Device syncs before TRT free were skipped — restore lm-trt.h:702-709 ordering (Golden rule 5). |
| Node timing table shows huge Adapter/gap times that engine wall-clock contradicts | stdout pipe-buffering skew — trust engine Applied in / Wall clock lines. |
[TRT-WARN] Using default stream in enqueueV3() | A call path passed a null stream; current code creates dedicated streams (dit-trt.h:562) — find the regressing caller. |
| VAE segfault on RTX 50xx (sm_120) via ORT-TRT | Known Blackwell Myelin fusion bug. The known workaround (builder_optimization_level=1) cannot be expressed via the legacy V1 EP options used in vae-ort.h:78-98 — fixing requires migrating to the V2 EP API. |
Stream ERROR: streaming requires TRT (ONNX model) | stream_mode requested with a GGUF DiT selected — switch to an ONNX DiT. |
| Wrong/stale engine after switching DiT models | Static TRT context keys on the ONNX path (s_trt_onnx_path, pipeline-synth-ops.cpp:1230-1234) and rebuilds on change — start staleness debugging there. |
| ONNX DiT + cover mode fails on FSQ | FSQ weights aren't in ONNX dirs; loader falls back to a hardcoded safetensors-dir list in pipeline-synth.cpp; log WARNING: no safetensors DiT found for FSQ. |
| Back-to-back generations each pay ~7.5s DiT load | Default EVICT_STRICT frees the TRT engine after every generation — enable co-resident ("Keep DiT & VAE loaded"). |
adapter-trt.h, log [Adapter-TRT] Applied in 1777 ms).hot-step-sampler.h:593 is authoritative.IRefitter, which ORT does not expose. ORT+TRT-EP remains correct for refit-free models (VAE, encoders).kv_cache_block_offsets). A built trtllm-engine-RTX5090/ still sits in models/onnx/lm-4B/ — it is inert.models/onnx/dit-fp8/ has no dit_fp8.engine on disk (checked 2026-07-02) — first native-TRT use would trigger a full rebuild; the dir may mainly serve as the home for the ONNX text/cond encoders.engine/docs/ARCHITECTURE.md — committed engine internals, CLI, request JSON.docs/dev/plugins-authoring.md — committed Lua plugin authoring (solvers/schedulers/guidance run identically on GGML and TRT samplers).docs/plans/2026-05-31-TRT-OPTIMIZATION-PLAN.md, docs/plans/2026-04-18-performance-optimizations.md, docs/plans/dit_optimization_analysis.md — local-only, gitignored; may be absent on other machines. Where they disagree with code, code wins (see the re-upload discrepancy above).© scragnog, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 1 other file in .claude/skills/engine-performance of scragnog/HOT-Step-CPP.
Open the folder on GitHubat commit 91e92a8
Engine Performance next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Engine Performance this skillscragnog/HOT-Step-CPP | 171 | — | ~4.9k | Automated safety check: Pass | MIT | |
| Matlab Use Visual Inspectionmatlab/matlab-agentic-toolkit | 1.1k | — | ~3.1k | Automated safety check: Pass | Custom licence | |
| Onboard Jetpack5 Inference BackendsEGalahad/sim2real | 145 | — | ~1.1k | Automated safety check: Pass | None | |
| Tao Finetune ClipNVIDIA/skills | 3.5k | — | ~4k | Automated safety check: Notes | Apache-2.0 | |
| Tao Finetune Video ClipNVIDIA/skills | 3.5k | — | ~3.5k | Automated safety check: Notes | Apache-2.0 | |
| Tao Port Huggingface ModelNVIDIA/skills | 3.5k | — | ~4.5k | Automated safety check: Notes | Apache-2.0 |
matlab/matlab-agentic-toolkit
Build machine vision inspection systems with MATLAB Visual Inspection Toolbox.
EGalahad/sim2real
Install, convert, debug, and benchmark sim2real ONNX GPU and TensorRT inference backends on onboard JetPack 5 Orin hosts such as g1-cable.
NVIDIA/skills
CLIP vision-language model for image-text retrieval, zero-shot classification, embedding extraction, ONNX export, and TensorRT deployment.
NVIDIA/skills
InternVideo2-CLIP L14 (TAO videoclip) for video-text retrieval, zero-shot classification, embedding extraction, LoRA fine-tuning, ONNX export, and TensorRT deployment.
NVIDIA/skills
Integrate a HuggingFace Computer Vision model into the NVIDIA TAO Toolkit ecosystem (tao-core config, tao-pytorch trainer, tao-deploy TensorRT pipeline).
majiayu000/spellbook
优化实际模型推理链路,将正确性对齐、分段 profiling、显存与数据搬运、TensorRT/ONNX/PyTorch 后端、attention/kernel、FP8/compile、缓存与少步采样、质量回归、GPU 成本和服务验收串成同一实验闭环。当用户要求推理提速、降低显存或 GPU 成本、复现模型效果、定位 GPU 利用率低、优化图像/视频/扩散模型或自托管 LLM 时使用,提供…
scragnog/HOT-Step-CPP
The standard way to run a listening test in HOT-Step - a local HTML score sheet next to the renders where Rob plays each track, scores it 1-5 on named criteria, and the page charts the two score…
scragnog/HOT-Step-CPP
Maps HOT-Step's native MiniMax-Music3 backend — engine port modules, endpoints, server/UI integration, parity/fixture infrastructure, and the hard-won trap list.
scragnog/HOT-Step-CPP
The validated recipe for training MiniMax-Music3 planner-LM style adapters (artist/album clones) with ace-train mm3-lm-train and the Training Studio.
scragnog/HOT-Step-CPP
Runbook for cutting and publishing a HOT-Step CPP release via a v git tag that triggers the multi-platform CI build and drafts a GitHub Release.
scragnog/HOT-Step-CPP
Safely pulls upstream acestep.cpp changes into the HOT-Step engine fork without destroying its integration hooks.
scragnog/HOT-Step-CPP
Diagnoses HOT-Step CPP generation failures, engine crashes, hangs, and startup problems from the logs/ session folders.
Works with
Categories
Explains where HOT-Step generation time goes (LM/DiT/VAE), how the TensorRT paths activate, how to benchmark from logs, and which knobs trade quality for speed. Engine Performance is an agent skill from scragnog/HOT-Step-CPP. Explains where HOT-Step generation time goes (LM/DiT/VAE), how the TensorRT paths activate, how to benchmark from logs, and which knobs trade quality for speed.
Engine Performance fits situations like: profiling slow generations; working on dit-trt.h/lm-trt.h/adapter-trt.h/hot-step-sampler-trt.h; debugging TRT/ONNX issues; deciding what performance work is DONE vs PLANNED.
Run `npx skills add scragnog/HOT-Step-CPP --skill engine-performance -a claude-code`. Or copy the skill folder (.claude/skills/engine-performance in scragnog/HOT-Step-CPP) into .claude/skills/engine-performance in your project. Claude Code loads it when a task matches its description.
Run `npx skills add scragnog/HOT-Step-CPP --skill engine-performance -a codex`. Or copy the skill folder (.claude/skills/engine-performance in scragnog/HOT-Step-CPP) into .agents/skills/engine-performance in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add scragnog/HOT-Step-CPP --skill engine-performance -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/engine-performance, .gemini/skills/engine-performance, .github/skills/engine-performance and .opencode/skills/engine-performance in your project.
Going by SKILL.md and its folder, Engine Performance needs the command-line tools its instructions call (cmake). Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Engine Performance is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.9k tokens (SKILL.md is roughly 20k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Engine Performance: Matlab Use Visual Inspection (matlab/matlab-agentic-toolkit, 1.1k stars), Onboard Jetpack5 Inference Backends (EGalahad/sim2real, 145 stars), Tao Finetune Clip (NVIDIA/skills, 3.5k stars) and Tao Finetune Video Clip (NVIDIA/skills, 3.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
scragnog (a GitHub user) maintains it in scragnog/HOT-Step-CPP, which has 171 GitHub stars. The repository holds 18 skills in this directory. The repository was last updated on October 7, 2026.
Source: scragnog/HOT-Step-CPP on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.