Aoti Debug
pytorch/pytorch
Debug AOTInductor (AOTI) errors and crashes. An agent skill from pytorch/pytorch.
How HOT-Step's custom flash-attention training ops (GGMLOPFLASHATTNTRAIN/BACK) work, what the AS1.5 DiT trainer campaign proved and disproved, and the exact contract for porting flash mode to the…
$ npx skills add scragnog/HOT-Step-CPP --skill flash-attn-training -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install scragnog/HOT-Step-CPP flash-attn-training --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/scragnog/HOT-Step-CPP.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/flash-attn-training .claude/skills/flash-attn-training && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "flash-attn-training" agent skill from https://github.com/scragnog/HOT-Step-CPP/tree/master/.claude/skills/flash-attn-training into .claude/skills/flash-attn-training/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "flash-attn-training", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/scragnog/HOT-Step-CPP/tree/master/.claude/skills/flash-attn-trainingType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add scragnog/HOT-Step-CPP --skill flash-attn-training -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install scragnog/HOT-Step-CPP flash-attn-training --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/scragnog/HOT-Step-CPP.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills/flash-attn-training .agents/skills/flash-attn-training && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "flash-attn-training" agent skill from https://github.com/scragnog/HOT-Step-CPP/tree/master/.claude/skills/flash-attn-training into .agents/skills/flash-attn-training/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "flash-attn-training", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add scragnog/HOT-Step-CPP --skill flash-attn-training -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install scragnog/HOT-Step-CPP flash-attn-training --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/scragnog/HOT-Step-CPP.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills/flash-attn-training .cursor/skills/flash-attn-training && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "flash-attn-training" agent skill from https://github.com/scragnog/HOT-Step-CPP/tree/master/.claude/skills/flash-attn-training into .cursor/skills/flash-attn-training/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "flash-attn-training", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/scragnog/HOT-Step-CPP.git --path .claude/skills/flash-attn-training--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add scragnog/HOT-Step-CPP --skill flash-attn-training -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install scragnog/HOT-Step-CPP flash-attn-training --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/scragnog/HOT-Step-CPP.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills/flash-attn-training .gemini/skills/flash-attn-training && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "flash-attn-training" agent skill from https://github.com/scragnog/HOT-Step-CPP/tree/master/.claude/skills/flash-attn-training into .gemini/skills/flash-attn-training/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "flash-attn-training", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install scragnog/HOT-Step-CPP flash-attn-trainingInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add scragnog/HOT-Step-CPP --skill flash-attn-training -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/scragnog/HOT-Step-CPP.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills/flash-attn-training .github/skills/flash-attn-training && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "flash-attn-training" agent skill from https://github.com/scragnog/HOT-Step-CPP/tree/master/.claude/skills/flash-attn-training into .github/skills/flash-attn-training/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "flash-attn-training", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add scragnog/HOT-Step-CPP --skill flash-attn-training -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install scragnog/HOT-Step-CPP flash-attn-training --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/scragnog/HOT-Step-CPP.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills/flash-attn-training .opencode/skills/flash-attn-training && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "flash-attn-training" agent skill from https://github.com/scragnog/HOT-Step-CPP/tree/master/.claude/skills/flash-attn-training into .opencode/skills/flash-attn-training/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "flash-attn-training", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
flash-attn-trainingHow HOT-Step's custom flash-attention training ops (GGMLOPFLASHATTNTRAIN/BACK) work, what the AS1.5 DiT trainer campaign proved and disproved, and the exact contract for porting flash mode to the…
Flash Attn Training is an agent skill from scragnog/HOT-Step-CPP. How HOT-Step's custom flash-attention training ops (GGMLOPFLASHATTNTRAIN/BACK) work, what the AS1.5 DiT trainer campaign proved and disproved, and the exact contract for porting flash mode to the other trainers (AS1.5 LM, MM3 LM, MM3 DiT). Use when adding --attn flash to any ace-train subcommand, touching engine/ggml/src/ggml-cuda/fattn-train., changing a trainer's VRAM model, debugging "flash is slower/uses more VRAM than expected", or interpreting any flash-vs-exact measurement.
Its SKILL.md is about 5.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Development. It works with CUDA. The repository describes itself as: Turn dials. Summon bangers! NOW WITH MORE C++! Local AI music generation powered by GGML. The licence is MIT.
8 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit eeeded6. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
gitFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Flash Attn Training loads about 5.4k tokens when it runs. Until then it costs about 128 tokens; SKILL.md has 2,969 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from scragnog/HOT-Step-CPP at commit eeeded6, republished under its MIT licence (© scragnog). 2,969 words, ~5,405 tokens.
.claude/skills/flash-attn-training/SKILL.md (or your agent's skills folder).Written 2026-09-02 from the AS1.5 DiT campaign (commits 28ca16d3 → 10c37556).
Everything here was measured on an RTX 5090 (32 GB, sm_120) unless it says otherwise.
The deep docs are gitignored, local-only (docs/plans/2026-09-01-flash-attn-backward.md,
fattn-train-spec.md, fattn-train-tf32-design.md); this skill is the committed distillation.
Context for a reader with zero prior exposure: ggml's autodiff had no attention backward,
so every trainer built attention as mul_mat → soft_max_ext → mul_mat and retained the
[S,S,Nh] softmax per layer for the backward — the O(S²) term that capped DiT training crops
at ~50 s of audio on 32 GB. We wrote our own fused forward+backward ops (CPU reference + CUDA
TF32 kernels), carried as commits on the HOT-ggml fork that engine/ggml pins (see docs/dev/ggml-fork.md). Attention memory is now
linear in S; the DiT auto-fit picks full-song crops. Rob ear-validated the first flash-trained
adapter as "fantastic".
| Piece | Where | Notes |
|---|---|---|
Ops GGML_OP_FLASH_ATTN_TRAIN / _BACK | engine/ggml/include/ggml.h, src/ggml.c (constructors, view getters, autodiff case), src/ggml-cpu/ops.cpp (f32 reference), src/ggml-backend-meta.cpp | Appended at the END of the op enum. Forward output is ONE packed tensor: O [D,Nh,S,B] then LSE [Nh,S,B]; ggml_flash_attn_train_get_o() views O out. Backward packs dQ|dK|dV. |
| CUDA kernels | engine/ggml/src/ggml-cuda/fattn-train.cu/.cuh (NEW files — never touch the inference fattn-*.cu/.cuh) | Scalar f32 v1 kernels kept as strict mode + pre-sm_80 fallback; TF32 mma (m16n8k8) kernels are the default. Bitwise-deterministic in every mode: no fp atomics, fixed schedules. |
| Precision knob | ggml_flash_attn_train_set_prec/get_prec (op_params slot 3) | GGML_PREC_DEFAULT (= 0 = zero-init!) → TF32 on sm_80+; GGML_PREC_F32 → v1 scalar. Autodiff copies the forward's prec onto the backward node. |
| Where it lives | HOT-ggml hot-step-neutral commits flash-attn-train (+ alloc-free-blocks), pinned by engine/ggml (docs/dev/ggml-fork.md) | Change the kernels on the fork, then move the gitlink. verify-hooks.ps1 Hook 12/13 grep the markers; Hook 17 refuses a dirty engine/ggml. |
| DiT trainer surface | ace-train train-dit --attn exact|flash|flash-f32 (default exact in the CLI; the Training Studio form defaults to flash) | dit_attn_flash() in engine/src/train/dit-train-graph.h beside the untouched dit_attn_f32(). Both self- and cross-attention route through it. |
| Parity harness | engine/tools/fattn-train-test.cpp, target fattn-train-test | --backend cpu|cuda|vulkan (default cpu), resolved by registry device name — never "first GPU found", and a missing requested backend is a hard failure, never a silent CPU fallback. --prec f32|tf32 (tf32 is CUDA-only; Vulkan runs f32 at 1e-4 same as CPU), --extra, --large, --bench, --bench-tr (CUDA-only; rejects --backend vulkan). |
| Profilers | --profile-step N (coarse buckets); DIT_PROFILE_NODES=1 per-node with site attribution (engine/src/train/dit-node-profile.h) | Node profiler is env-gated, zero cost when off. |
| Server/UI | attnBackend: 'exact'|'flash'|'flash-f32' through types.ts → routes/training.ts → aceTrain.ts; Training Studio checkbox | cropMax 0 = "no pin" end to end (see trap 6). |
Every trainer that gains flash mode must keep all of these. Each one exists because its absence bit us.
exact, and exact means byte-identical. With the flag
off the emitted graph must be the pre-flash graph to the byte — the DiT proves it with T3
(0.00e+00 on 17 named taps) and SC1–SC3 (0.000e+00 grad delta) against a
reverted-tree baseline. Gate the mode at the attention call sites only; restructure nothing else.supports_op probe at trainer init, hard error on false. ggml_backend_supports_op
returning false is NOT an error in this engine: backend_sched_new registers the CPU backend
alongside CUDA, so the scheduler silently splits attention onto the CPU — correct, unusably
slow, low VRAM, tripwire silent, i.e. indistinguishable from a pass on every number the run
reports. Build a scratch no_alloc node pair at the run's REAL shapes (both attention sites,
effective Nkv) and abort with a named error. See DiT dit-train-run.h "spec 9.8 probe".--bwd mm. DiT: over 200 same-seed epochs flash
drifted less than --bwd mm. Not identity — never claim identity.attn_prec) in the run log, not just the requested mode.
Reason: op_params zero-init == GGML_PREC_DEFAULT, so every --attn flash run on Ampere+
was ALREADY TF32 before the knob existed and said nothing about it.Adoption is call-site wiring, not kernel work. The ops take any additive F16 mask
([S_kv,S] or [S_kv,S,1,B] broadcast), GQA (Nkv < Nh at B=1), S_kv ≠ S, and
non-contiguous q/k/v views (only nb[0]==4 is required — do not ggml_cont them, that
gives back the VRAM win).
| Trainer | Files | Specifics |
|---|---|---|
| AS1.5 LM (R2) — DONE 2026-09-02 | engine/src/train/lm-graph.h, lm-train-run.h, lm-vram.h, lm-selftest.h, flash-prec.h | Causal = one triangular −INF mask; the kernel skips all-−INF tiles, so causal gets ~half its compute skipped free. Qwen GQA at B=1 is the tested path. Ships off by default (CLI and Training Studio checkbox); 4B low-VRAM is 5.5% faster than the shipped head-blocked arm and 1.2% faster at equal graph shape (opposite split from the DiT — there the fused kernel is the whole win, here the head-block copies are); naive 0.6B roughly doubles auto-fit maxLen, 1.7B only 1.27×. Not yet ear-validated — see project-flash-attn-backward.md in memory and §7/§8 below for the full numbers and open items. |
| MM3 LM (R3) — DONE 2026-09-05 | mm3-lm-train-run.h, mm3-lm-adapter.h, mm3-lm-graph.h | The "sequence term was quadratic all along" retained softmax is what goes — but refused rather than composed with a frozen/trained KV prefix: --attn flash is rejected together with --prefix-frames > 0 or --prefix-n > 0 (the fused kernel doesn't take the rectangular mask a prefix needs), so the no-dK/dV-for-frozen-columns idea above was never built. Default exact. Measured (RTX 5090, mm3-lm-f16/mm3-lm-q8_0, oasis_morningglory, rank 256, checkpointed): flash is within noise of exact up to ~1500 frames (checkpointing already hides the small softmax in allocator slack), then saves VRAM growing to ~9 GB by 5000 frames; the usable crop ceiling moves from ~4300 frames (exact, before it starts spilling past ~29 GB used) to at least 11,178 frames (flash — this corpus's longest track, no OOM reached). Paired 20-step run at the shipped recipe's crop (750): 2118 ms/step flash vs 2215 ms exact, max loss drift 1.9e-4. Resolves to tf32 on this card for --attn flash, f32 for --attn flash-f32. Not ear-validated — the shipped recipe still trains at crop 750, where flash measures no benefit. Full numbers: docs/dev/training-internals.md MM3 section. |
| MM3 DiT (R4) | mm3-dit-train-*.h | Bidirectional like the AS DiT; smallest win (shorter sequences). |
| YuE2 joint (AITK) — already flash, always | engine/src/train/yue2-aitk-graph.h block() | Every AR and NAR attention call is ggml_flash_attn_train (TF32; YUE2_AITK_STRICT_F32 = F32). The quadratic chain exists only under YUE2_AITK_DIAGNOSTIC_MATH_ATTENTION. The log's attention_forward: "tf32" is this kernel. No port needed; the decoder window is --nar-crop-frames (2026-09-24). Don't mistake it for an exact-attention trainer (an agent did). |
For each: (a) sibling xxx_attn_flash() returning exactly the shape the manual chain returned;
(b) flag + log fields; (c) probe; (d) selftest rung; (e) VRAM branch; (f) drift A/B;
(g) --bench-tr-style measurement at that trainer's REAL geometries (see §4).
--bench (window mask only) flattered fused. The trainer has
three: windowed self (fused 0.89× cuBLAS), full self with NO mask (1.18×), cross at
S_kv = enc_S (1.52×). --bench-tr covers all three. Half the DiT layers are full attention
(layer_type = i % 2) and get no tile skip.DIT_PROFILE_NODES=1 found the whole flash deficit is the
cross-attention BACKWARD (67.8 vs 34.8 ms/step); self-attention is a wash, cross forward is
2× faster. Root cause: both TF32 backward kernels split warps by output d-range and recompute
the shared S/dP tiles (dK/dV 1.5×, dQ 1.67× the needed mma). A dQ role split measured −2.7%
end to end → reverted under a 3% bar. dK/dV split is blocked by ~128 B of static shared
memory at the 3-blocks/SM occupancy cliff. Recorded in the plan doc; not worked around.dit_vram_arena_bytes_flash) takes enc_S explicitly; cross-attention scales
with enc_S×S and at crop 1250 exceeds self-attention's S² — "enc_S is small" was refuted.DIT_FLASH_LOKR_RETENTION (0.62) was fitted before the LoKR
apply reorder and now over-predicts +16.5% (safe direction, ~one crop step unspent). Owed
refit; the batch>1 term is B=1-fitted and over-conservative.crop_max to the dataset's longest track ONLY when the user passed
no --crop-max; a.crop_max_user is the pin flag. See trap 6..cu files need a cmake re-configure (the ggml-cuda CMake globs *.cu).ace-server holds ggml-base.dll/ggml-cuda.dll; any ggml change
needs the app down (/api/shutdown or dev-rebuild.bat). ace-train.exe is NOT held, so
trainer-only edits build with the app up. Never kill ace-server (Node respawns it).ggml_scale(packed, 0) + ggml_acc(dO); garbage in the O→LSE alignment gap becomes NaN.
Both CUDA and CPU forwards zero the gap explicitly. Zero-width at every tested geometry, so
tests never see it — keep the memset.ggml_scale is in ggml_op_can_inplace; it is safe only because
the packed tensor always has a view child. The backward asserts dst->data != fwd->data.dit_expand_heads exists. Flash mode skips the
expansion (native GQA), which also disarms the CUDA REPEAT_BACK cap on Nkv·max(S,enc_S)·B.
Measured: batch 1 still wins on throughput and loss.--crop-max, which the engine treats as a user pin → the flash
lift never fired from the UI. cropMax 0 now means "omit the flag". Quality presets must not
re-pin it in flash mode. Any new trainer flag with an engine-side "user set it" sentinel has
this exact failure mode — check the arg emitter.ggml_set_loss only allocates)
and assert a non-zero reference gradient, or both arms compare 0 vs 0 and pass vacuously.dit_sa_mask never produces a fully-masked key column (pad columns stay open for padded
query rows) — use dit_ca_mask for the exactly-zero-gradient assertion.soft_max_ext produces NaN.
Exclude them from reference diffs, check them directly.tile<16,8,float> is the C/D map;
using it as the tf32 A operand gives deterministic garbage. Derive with a probe kernel.core.autocrlf=true is CRLF and every
hunk fails. Replay with git -c core.autocrlf=false -c core.eol=lf archive. Export patches
hunk-filtered: several patches share ggml.c and ggml-cuda.cu.MAX_FREE_BLOCKS (ggml-alloc) was 256; LoKR dim 256 (19k-node graph) overflowed it.
Now 1024 via alloc-free-blocks.patch. Inference-shared → smoke generation after touching.--mirror bf16 means bf16 COMPUTE, not just bf16 storage — it rounds activations and
gradients at every trainable-layer GEMM, and the adapters it trains are audibly coarse
("bitty", Rob 2026-09-02). Use --mirror bf16-f32: same BF16 residency, an in-graph
ggml_cast to F32 at each mul_mat site, and over 12 same-seed epochs on mika it is
bit-identical to --mirror f32 while bf16 drifts to 7.8e-3. It costs ~180 MB of transient
arena and ~25% step time against f32 at equal crop, and buys 2.5× the flash auto-fit crop
(1542 vs 610). Only --bwd mm carries it — the out_prod fallback arm keeps the forward
cast alive and silently spends the ~8 GB back.fattn-train-test --bench-lm's blocked arm fed a ggml_cont(view) straight into the
reference attention chain, whose backward hands back a transposed (non-contiguous) gradient
— GGML_OP_CONT's backward asserts on that and the tool produced no table at all. The
trainer never hits it because a ggml_reshape always sits between the cont and the chain,
and RESHAPE's backward re-conts. Fix: wrap each bench-arm tensor in a shape-preserving
ggml_reshape too, so the bench pays the same backward copy the trainer pays. Any bench
harness that hand-builds a reference graph needs to mirror the trainer's node shapes, not
just its op sequence.--max-len filters, it does not truncate. Songs longer than it are skipped outright, so
alloc_seq = min(max_len, longest SURVIVING sample) — pinning a value above the whole
corpus's longest track yields an empty dataset (no-samples), and a VRAM-model cell "at
S=1024" is really whatever the longest surviving song happens to be. Pick the dataset for
the S you want, then report the actual S; don't trust the flag to hit a number.maxLen whose own estMb already exceeds free
VRAM, then die on cudaMalloc with a hard access violation (0xC0000005) instead of a
clean lm_fatal — reproduces identically on a pre-flash binary, so it is not new. Root
cause is the same non-attention polynomial (c2f/c2h) the flash branch's
naive_nonattn_scale now corrects around; the exact-mode fix is owed (see §8) and needs its
own gate since it moves every shipped run's estMb.| Measurement | Value |
|---|---|
| Fused TF32 vs cuBLAS per site, fwd+bwd, window mask | 0.94× / 0.64× / 0.49× at S=625/1250/3000 |
| Same at the trainer's real geometries | windowed 0.89×, full-self 1.18×, cross(S_kv 1877) 1.52× |
| Attention VRAM per site at S=3000 | 487 MB fused vs 4.9 GB manual |
| Parity worst rel err | f32 3.5e-6 (bar 1e-4); tf32 4.7e-4 (bar 5e-3, floor 1e-5) |
| Flash vs exact drift, 200 same-seed epochs | smaller than --bwd mm |
| Done-gate auto-fit, production LoKR, unpinned | albumJ 1498 (enc_S 1877), album D 1616 (enc_S 640); LoRA r16 ~3400 |
| LoKR apply reorder | −10% step, LoKR:LoRA 1.35→1.21; the two copies are unavoidable, ~7% of step |
| 12 GB emulated card, flash+bf16+LoRA r16 | full 32-layer depth, crop 410, 4 segments |
LM, 4B low-VRAM, flash vs shipped (exact --attn-head-block 8) | 5.5% faster/micro-step, 3.8% lower peak VRAM (paired, interleaved, albumF substitute) |
LM, 4B low-VRAM, flash vs equal-shape (exact --attn-head-block 0) | 1.2% faster — the head-block copies are almost the whole DiT-vs-LM difference |
LM attention-only bound (fattn-train-test --bench-lm vs blocked) | 0.74×/0.79×/0.80× at S=1024/2113/3500 |
LM naive auto-fit maxLen lift, flash vs exact | 0.6B ~2.0× (3136→6208 tok); 1.7B ~1.27× (2624→3328 tok) |
| LM 50-epoch same-seed drift, flash vs exact | same class as --weights bf16; smaller on 2/3 measures, ~20% larger on final CE (1 seed, no error bar) |
| MM3 LM, usable crop ceiling, flash vs exact | ~4300 frames exact -> >=11,178 frames flash (this corpus's longest track; RTX 5090, oasis_morningglory, rank 256) |
| MM3 LM, paired step time at the shipped recipe's crop (750 frames) | 2118 ms/step flash vs 2215 ms exact |
DIT_FLASH_LOKR_RETENTION refit after the apply reorder; batch>1 VRAM term.c2f/c2h non-attention polynomial is ~2.2× light on the naive path
(−11.9% to −12.9% measured, same class as the DiT's exact-mode item above); the flash branch's
naive_nonattn_scale corrects around it but the exact-mode fix itself is owed and needs its
own gate, since it would move every shipped run's estMb/auto-fit maxLen.albumF, not album I — the box has no albumI* tensor dir, and
the plan's ear pair (G7) is specified on album I/E3 lineage. album I codes need Preprocess +
Extract via the Training Studio batch pipeline before G7 can run as written._experiments/_LISTENING) — not run, needs
Rob; the flash checkbox stays off until it lands.mm3-lm-train crashes at export with a ggml-backend.cpp tensor-write-out-of-bounds
assert; the LM exact-mode naive auto-fit can pick a maxLen that OOMs via access violation
instead of a clean fatal (trap 19).© scragnog, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .claude/skills/flash-attn-training of scragnog/HOT-Step-CPP.
Open the folder on GitHubat commit eeeded6
Flash Attn Training next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Flash Attn Training this skillscragnog/HOT-Step-CPP | 174 | — | ~5.4k | Automated safety check: Pass | MIT | |
| Aoti Debugpytorch/pytorch | 104k | 1 repos | ~1.7k | Automated safety check: Pass | Custom licence | |
| Debug Distributed Hangsgl-project/sglang | 37k | 2 repos | ~2.4k | Automated safety check: Pass | Apache-2.0 | |
| Create Cuda Python Pull RequestNVIDIA/cuda-python | 3.4k | — | ~1.1k | Automated safety check: Pass | Apache-2.0 | |
| CUTLASS FMHA Incremental Rebuildmicrosoft/onnxruntime | 22k | — | ~1.3k | Automated safety check: Pass | MIT | |
| ONNX Runtime Source Buildmicrosoft/onnxruntime | 22k | — | ~1.4k | Automated safety check: Pass | MIT |
pytorch/pytorch
Debug AOTInductor (AOTI) errors and crashes. An agent skill from pytorch/pytorch.
sgl-project/sglang
Debug hanging issues in SGLang distributed inference (TP/PP/DP/EP).
NVIDIA/cuda-python
Create a CUDA Python pull request from an approved personal or organization-owned fork, including the GitHub CLI GraphQL fallback for renamed organization-owned forks.
microsoft/onnxruntime
Explains why editing CUTLASS fused-MHA headers in ONNX Runtime can leave stale CUDA kernels after an incremental build, and how to force and verify a real rebuild.
microsoft/onnxruntime
Builds ONNX Runtime from source with its build scripts, explaining the update, build and test phases, key flags and where the build output lands.
stas00/the-art-of-debugging
Condensed debugging method and tool recipes for Unix, Python and PyTorch programs: crashes, hangs, segfaults, wrong output, CUDA OOM, NaN values and slowness.
scragnog/HOT-Step-CPP
The standard way to run a listening test in HOT-Step - a local HTML score sheet next to the renders where Rob plays each track, scores it 1-5 on named criteria, and the page charts the two score…
scragnog/HOT-Step-CPP
Explains where HOT-Step generation time goes (LM/DiT/VAE), how the TensorRT paths activate, how to benchmark from logs, and which knobs trade quality for speed.
scragnog/HOT-Step-CPP
Maps HOT-Step's native MiniMax-Music3 backend — engine port modules, endpoints, server/UI integration, parity/fixture infrastructure, and the hard-won trap list.
scragnog/HOT-Step-CPP
The validated recipe for training MiniMax-Music3 planner-LM style adapters (artist/album clones) with ace-train mm3-lm-train and the Training Studio.
scragnog/HOT-Step-CPP
Runbook for cutting and publishing a HOT-Step CPP release via a v git tag that triggers the multi-platform CI build and drafts a GitHub Release.
scragnog/HOT-Step-CPP
Safely pulls upstream acestep.cpp changes into the HOT-Step engine fork without destroying its integration hooks.
Works with
Categories
How HOT-Step's custom flash-attention training ops (GGMLOPFLASHATTNTRAIN/BACK) work, what the AS1.5 DiT trainer campaign proved and disproved, and the exact contract for porting flash mode to the…. Flash Attn Training is an agent skill from scragnog/HOT-Step-CPP.5 LM, MM3 LM, MM3 DiT).
Flash Attn Training fits situations like: adding --attn flash to any ace-train subcommand; touching engine/ggml/src/ggml-cuda/fattn-train; changing a trainers VRAM model; debugging flash is slower/uses more VRAM than expected.
Run `npx skills add scragnog/HOT-Step-CPP --skill flash-attn-training -a claude-code`. Or copy the skill folder (.claude/skills/flash-attn-training in scragnog/HOT-Step-CPP) into .claude/skills/flash-attn-training in your project. Claude Code loads it when a task matches its description.
Run `npx skills add scragnog/HOT-Step-CPP --skill flash-attn-training -a codex`. Or copy the skill folder (.claude/skills/flash-attn-training in scragnog/HOT-Step-CPP) into .agents/skills/flash-attn-training in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add scragnog/HOT-Step-CPP --skill flash-attn-training -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/flash-attn-training, .gemini/skills/flash-attn-training, .github/skills/flash-attn-training and .opencode/skills/flash-attn-training in your project.
Going by SKILL.md and its folder, Flash Attn Training needs the command-line tools its instructions call (git).
SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Flash Attn Training is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 5.4k tokens (SKILL.md is roughly 22k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Flash Attn Training: Aoti Debug (pytorch/pytorch, 104k stars), Debug Distributed Hang (sgl-project/sglang, 37k stars), Create Cuda Python Pull Request (NVIDIA/cuda-python, 3.4k stars) and CUTLASS FMHA Incremental Rebuild (microsoft/onnxruntime, 22k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
scragnog (a GitHub user) maintains it in scragnog/HOT-Step-CPP, which has 174 GitHub stars. The repository holds 18 skills in this directory. The repository was last updated on October 9, 2026.
Source: scragnog/HOT-Step-CPP on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.