Agent skill

Adapter System

by scragnog in scragnog/HOT-Step-CPP

Explains how HOT-Step's LoRA/LoKr adapter system loads, merges, caches, stacks, and regionally masks adapters at runtime, including hard-won failure modes.

MITAuto-check passedAI & LLM Engineering

Install Adapter System

skills CLI
$ npx skills add scragnog/HOT-Step-CPP --skill adapter-system -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install scragnog/HOT-Step-CPP adapter-system --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/scragnog/HOT-Step-CPP.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/adapter-system .claude/skills/adapter-system && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
adapter-system
GitHub stars
173
Token cost
~6.3k tokens
SKILL.md length
2,946 words
Files
2
Skills in repo
18
Repo updated
First seen
Licence
MIT

At a glance

Explains how HOT-Step's LoRA/LoKr adapter system loads, merges, caches, stacks, and regionally masks adapters at runtime, including hard-won failure modes.

  • Works in 10 steps: GGML clobbers input-tensor buffers after… → Cache-key discipline (commit 168dcb5 —… → Basin re-base applies ONCE PER STACK,… → …
  • Working on adapter loading/merging
  • SKILL.md covers Glossary (read first — no…, When to use this skill, Golden rules (each one… and How an adapter reaches the GPU…, plus 7 more sections
  • Calls cmake, npx and npm

What it does

Adapter System is an agent skill from scragnog/HOT-Step-CPP. Explains how HOT-Step's LoRA/LoKr adapter system loads, merges, caches, stacks, and regionally masks adapters at runtime, including hard-won failure modes. Use when working on adapter loading/merging, multi-adapter stacking, per-section adapter masking, adapter cache keys, cross-base adapter conversion, debugging adapters that sound wrong or silent, or when UI adapter knobs (Adapter VRAM, Alignment Timing, Sum/Blend, trigger words, [Section]{k=v} lyric directives) appear dead or ignored.

Its SKILL.md is about 6.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `reference.md`).

It sits in AI & LLM Engineering, covering Fine-tuning. It works with C++. The repository describes itself as: Turn dials. Summon bangers! NOW WITH MORE C++! Local AI music generation powered by GGML. The licence is MIT.

When your agent uses it

  • Working on adapter loading/merging
  • Multi-adapter stacking
  • Per-section adapter masking
  • Adapter cache keys

Example prompts

  • “Use the adapter-system skill to explain how HOT-Step's LoRA/LoKr adapter system loads, merges, caches, stacks, and regionally masks adapters at…”
  • “/adapter-system”

Requirements

  • Node.js

Workflow steps

10 steps, taken from the first numbered list in SKILL.md.

  1. GGML clobbers input-tensor buffers after every compute. Any constant GPU input — especially the per-frame LoRA section masks — must be…
  2. Cache-key discipline (commit 168dcb5 — the law). Every input that changes the merged weights or loaded deltas MUST be part of ModelKey…
  3. Basin re-base applies ONCE PER STACK, first adapter only (merge: engine/src/dit.h:565-586; runtime: adapter_runtime_rebase in…
  4. Do NOT re-attempt regional self-attention isolation. Implemented in commit 0f3bf6d, reverted in ee041e1: penalising cross-section…
  5. Weight similarity does NOT imply adapter transferability. LoKR cross-base conversion to non-turbo XL bases fails even though the bases are…
  6. Never use real Q4_K for runtime delta storage. Its per-superblock optimizer stalls the load for minutes across ~360 tensors. The "q4_k"…
  7. Never assume runtime deltas are BF16. They may be Q4_0/Q8_0. mul_mat dequantizes transparently, but any code that reads delta bytes…
  8. LM echo sideband gotcha: never whitelist synth-request fields. Sideband/ServerFields-only params (adapter_runtime_quant…
  9. Edit the right files. The compiled server is engine/tools/hot-step-server.cpp — engine/tools/ace-server.cpp is upstream reference code…
  10. Build & git rules. After ANY engine edit: .\dev-rebuild.bat from repo root, immediately — NEVER engine/build.cmd directly (you cannot…

What it can do on your machine

Read from SKILL.md and the folder at commit 91e92a8. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • cmake
    • npx
    • npm
    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use npx, npm and git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Adapter System loads about 6.3k tokens when it runs. Until then it costs about 127 tokens; SKILL.md has 2,946 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~127
When it runs · the whole SKILL.md, loaded when a task matches
~6.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from scragnog/HOT-Step-CPP at commit 91e92a8, republished under its MIT licence (© scragnog). 2,946 words, ~6,254 tokens.

Download SKILL.mdSave it as .claude/skills/adapter-system/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
adapter-system
description
Explains how HOT-Step's LoRA/LoKr adapter system loads, merges, caches, stacks, and regionally masks adapters at runtime, including hard-won failure modes. Use when working on adapter loading/merging, multi-adapter stacking, per-section adapter masking, adapter cache keys, cross-base adapter conversion, debugging adapters that sound wrong or silent, or when UI adapter knobs (Adapter VRAM, Alignment Timing, Sum/Blend, trigger words, [Section]{k=v} lyric directives) appear dead or ignored.

Adapter (LoRA/LoKR) System

The deepest institutional-knowledge subsystem in HOT-Step CPP. Every rule below was paid for with a shipped bug or a failed research direction. Deeper detail (merge/runtime internals, cache-key construction, cross-arch research history) lives in reference.md in this folder.

Glossary (read first — no prior context assumed)

  • DiT — Diffusion Transformer, the C++ model that denoises audio latents (32 layers on the XL base). Lives in engine/src/dit.h / dit-graph.h.
  • LoRA — low-rank adapter: two small matrices A, B whose product B@A is a delta added to a base weight. LoKr — LyCORIS Kronecker-product variant (.lokr_w1/.lokr_w2 tensors). DoRA — adds a per-row multiplicative rescale; two on-disk namings, both supported in MERGE MODE ONLY: LyCORIS dora_scale (LoKr path) and PEFT lora_magnitude_vector (LoRA path, since 2026-07-19). With DoRA the delta carries only alpha/rank; user strength × group scale blends the decompose factor (s = u·(m/‖W+Δ‖) + (1−u)), so strength 0 does NOT fully disable a DoRA adapter.
  • Delta (Δ) — the full-size weight difference an adapter contributes: W_adapted = W_base + scale·Δ.
  • Merge mode — deltas are baked into base weights on load, before GPU upload (adapter-merge.h). Runtime mode — deltas live as separate GPU tensors, applied each sampling step as y = W@x + Δ@x (adapter-runtime.h + dit-graph.h).
  • aceReq — the JSON generation request (AceRequest) the Node server sends to the C++ engine's /synth endpoint.
  • Sideband — extra HOT-Step-only request fields the engine parses separately from AceRequest (struct ServerFields, engine/tools/hot-step-server.cpp:569) and stores in the global g_hotstep_params (engine/src/hot-step-params.h:205).
  • ModelKey — the cache key under which a loaded DiT is stored in model-store.{h,cpp}. Two requests with the same key reuse the same loaded (adapter-merged) model.
  • Basin re-base — before merging, nudge the loaded base T toward the adapter's original training base S: base ← base + β·(S − base) (see Institutional knowledge).

When to use this skill

  • Adding/changing anything in engine/src/adapter-*.h, the adapter parts of dit.h, dit-graph.h, hot-step-sampler.h, or hot-step-server.cpp.
  • Working on multi-adapter stacking, Sum/Blend, per-section [Header]{k=v} lyric directives, adapter VRAM quantization, or trigger words.
  • Debugging: adapter inaudible, wrong strength, wrong flavor after toggling settings, CUDA errors during sampling with section adapters, slow adapter loads.
  • Anything involving converting adapters between DiT bases or architectures.

Golden rules (each one reverses a shipped bug or dead research month)

  1. GGML clobbers input-tensor buffers after every compute. Any constant GPU input — especially the per-frame LoRA section masks — must be re-uploaded every sampling step, in all compute paths, and again after any mid-sampling graph rebuild (CFG-cutoff). Upload-once produces the "nil adapter" bug: masks read ~0, output is pure base model. Existing re-upload sites: engine/src/hot-step-sampler.h:605, :659, :716, :940, :1097. Any NEW graph input must join every one of those sites. (Commits a6db135, 2052bc3.)
  2. Cache-key discipline (commit 168dcb5 — the law). Every input that changes the merged weights or loaded deltas MUST be part of ModelKey, and a failed adapter load must NEVER be cached as a success. Key construction: engine/src/pipeline-synth.cpp:168-199; hash/eq: engine/src/model-store.cpp:48-94; failure path: engine/src/dit.h:648-658. If you add any parameter that affects merged/loaded weights, add it to the key AND the hash AND the equality — all three.
  3. Basin re-base applies ONCE PER STACK, first adapter only (merge: engine/src/dit.h:565-586; runtime: adapter_runtime_rebase in adapter-runtime.h, invoked after the first adapter stages). Applying it per adapter resets the running base on every merge (at β=1 only the last adapter survives) — in runtime mode it would duplicate the base correction N×. Works in BOTH modes since 7aac3c2 (runtime folds β·(S−T) into the staged delta sum — matmul distributivity makes it output-identical to merge). NOT supported on the per-section masking path (masked per-frame deltas can't carry an always-on base correction; engine warns + skips, and the cache key mirrors that skip in pipeline-synth.cpp).
  4. Do NOT re-attempt regional self-attention isolation. Implemented in commit 0f3bf6d, reverted in ee041e1: penalising cross-section self-attn logits broke musical continuity (multi-second silence gaps, degenerate later sections). Self-attention carries BOTH adapter identity AND musical coherence — you cannot suppress one without the other. The adapter_section_isolation param is still plumbed end-to-end but dormant (engine ignores it). See reference.md §Reverted for the structurally safer untried alternative.
  5. Weight similarity does NOT imply adapter transferability. LoKR cross-base conversion to non-turbo XL bases fails even though the bases are ~99% weight-identical — the root cause is basin-sensitivity, not weight drift (Institutional knowledge below).
  6. Never use real Q4_K for runtime delta storage. Its per-superblock optimizer stalls the load for minutes across ~360 tensors. The "q4_k" string is now aliased to Q4_0 (engine/src/adapter-runtime.h:136-141). Q4_0/Q8_0 require ne0 % 32 == 0; anything else falls back to BF16.
  7. Never assume runtime deltas are BF16. They may be Q4_0/Q8_0. mul_mat dequantizes transparently, but any code that reads delta bytes directly assuming BF16 crashes (a delta-L2 diagnostic did exactly this — removed in 46603bf).
  8. LM echo sideband gotcha: never whitelist synth-request fields. Sideband/ServerFields-only params (adapter_runtime_quant, adapter_section_align_at, rebase_*, …) do NOT survive the engine /lm round trip. server/src/routes/generate.ts rebuilds synth requests as {...aceReq, ...LM-generated-fields-only} (lines 238 and 320). A field whitelist here once made the Adapter VRAM and Alignment Timing knobs silently dead on the default path (fixed in 8ea519b/168dcb5).
  9. Edit the right files. The compiled server is engine/tools/hot-step-server.cpp — engine/tools/ace-server.cpp is upstream reference code, not compiled. The live sampler is engine/src/hot-step-sampler.h — dit-sampler.h is dead upstream code. pipeline-synth-ops.cpp:9 includes hot-step-sampler.h; losing that include on an upstream sync is silent (compiles, all solvers/schedulers/guidance go dead). After any sync run engine/verify-hooks.ps1.
  10. Build & git rules. After ANY engine edit: .\dev-rebuild.bat from repo root, immediately — NEVER engine/build.cmd directly (you cannot reliably tell whether the app is running; Node auto-respawns ace-server → infinite respawn + file-lock loop), never cmake --clean-first (20+ min CUDA recompile). TypeScript changes: npx tsc --noEmit, not npm run build. Git: stage explicit paths only — never git add -A or add -f (re-adds gitignored models/checkpoints/local docs); commit locally often; push only with explicit user approval.

How an adapter reaches the GPU (request flow)

  1. UI → Node: server/src/services/generation/translateParams.ts:83-143 maps params.loraStack → req.adapters [{name, scale}] (supersedes the single req.adapter), plus adapter_mode, adapter_runtime_quant, per-group scales, basin re-base fields (through :128), and trigger words for all stacked adapters (:131-143). Per-section directives in lyrics are parsed here (see below).
  2. Node → engine /synth: the engine parses the JSON twice — once as AceRequest (engine/src/request.{h,cpp}), once as the ServerFields sideband (hot-step-server.cpp:569+).
  3. Worker resolves and stages: adapter names → paths via the registry, with an absolute-path fallback for unregistered adapters (hot-step-server.cpp:1079). Fills g_hotstep_params.adapters and all sideband fields BEFORE ace_synth_load — the merge reads them during load (hot-step-server.cpp:1129 comment).
  4. Cache lookup: pipeline-synth.cpp:168-199 builds the DiT ModelKey; model-store returns the cached DiT or loads via dit_ggml_load.
  5. Load (dit.h): merge mode → sequential adapter_merge per stacked adapter (deltas accumulate: W ← W + s1·Δ1 + s2·Δ2 + …); runtime mode → skip QKV/gate_up fusion (dit.h:420 — runtime deltas need individual projections), then after GPU alloc load runtime deltas (dit.h:608+).
  6. Sampling (hot-step-sampler.h): deltas applied per step; section masks (if active) built and re-uploaded every step.

Per-section (regional) adapter masking — shipped, P1+P2

Different adapters active in different song sections, driven by lyric directives like [Chorus]{albumB2=1; albumN=0} (keys = adapter filename stem, or positional #2/2, 1-based).

  • Gate: activates only with directives in lyrics AND ≥2 stacked adapters (translateParams.ts:104). Forces adapter_mode=runtime (server :114; engine double-checks, hot-step-server.cpp:1157-1159). When the gate is unmet, directives are still stripped from lyrics (stripAdapterDirectives) so they never reach the LM/encoder as garbage tokens. The UI stack controls sit behind Advanced mode (advancedAdapters in ui/src/stores/globalParamsStore.ts).
  • Parser (server/src/services/generation/adapterSections.ts): Sum/Blend applied per section (blend budget default 0.75); directive-less sections use stack default scales; {…} with no key=val pair is treated as lyric text and left alone; all-typo keys warn and fall back to defaults; weights clamp ≥ 0.
  • Engine load (dit.h:618-640): each adapter goes into its OWN DiTLoRA in m->loras[], UNIT-scaled (1.0) — the per-frame mask carries the effective scale; loading with the stack scale would double-scale (dit.h:628-631). This costs N× VRAM vs the summed path, hence Q4_0 delta quantization (~¼ size).
  • Graph (dit-graph.h): one [1,S,1] F32 mask tensor per adapter, shared across layers. Frame-indexed projections get Σᵢ (Δᵢ@x) ⊙ maskᵢ; token/global projections (cross-attn k/v, cond_embed) get the adapter's scalar mean section weight instead — a frame mask is the wrong axis there (dit-graph.h:430-433). Known limitation: text conditioning is therefore a constant blend across the song.
  • Sampler (hot-step-sampler.h): P1 builds an initial frame→section map proportional to section character counts (:306-323), with a ~0.5 s triangular crossfade between sections (ce57675). P2: at adapter_section_align_at fraction of steps (default 0.55, UI "Alignment Timing"), estimate x0, run alignment extraction on a private scheduler, map frames → dominant lyric token → section via the header-anchored token map (pipeline-synth-ops.cpp:1419-1495), median-smooth, rebuild masks (:338-385). Falls back to the P1 map on failure.

Timestep-dependent adapter gating (interval experts / MoE) — shipped 2026-07-20

Per-adapter gain curves g(t) over flow-matching t (1=noise → 0=clean): each stacked adapter's per-frame mask is multiplied by its interpolated g(t) at EVERY model evaluation (upload_lora_masks(t_val) in hot-step-sampler.h — the single helper that replaced all five raw mask-upload sites). Different adapters can own different slices of the trajectory ("structure" expert early / "timbre" expert late — TD-LoRA / TimeStep Master scalar mixing). Key facts:

  • Rides the per-section machinery: a curve-carrying stack with no lyric directives gets a synthetic single whole-song section from translateParams.ts (+ forced runtime mode); allowed at stack size 1 (gates in dit.h and hot-step-server.cpp accept >=2 adapters OR gains active). Composes with real per-section directives (mask × gain).
  • Curves are NOT in the ModelKey — deliberate: they change per-step mask uploads only, never loaded weights, so mixing retunes with zero model reload (the |sect marker still applies via the synthetic section). Do not "fix" this by adding them to the key.
  • Direct-graph caveat: masks are pinned resident on direct graphs and normally skip per-step re-upload; with gains active upload_lora_masks runs on EVERY evaluation regardless (!dit_graph.direct || lora_gains_active). Host masks stay UNSCALED — scaling happens into a scratch buffer at upload so P2 rebuilds and repeated evals never compound gains.
  • Wire format: adapters[].gain_curve = uniform samples of g(x) over [0,1] (33 from the UI) + gain_domain: "steps" (x = remaining-steps fraction, engine maps each eval's t to its nearest schedule index) or "t" (x = raw flow-matching t). UI windows are ALWAYS "steps"; trained-expert / router curves must be "t" (must match the training axis). This split exists because shifted schedules are wildly nonuniform in t — shift-3 @ 20 steps spends 17 steps above t=0.5, so a t-domain "0–50%" window starved the late adapter to 3 tail steps ("all I hear is the first adapter"). Window edges are smoothstep ramps CENTERED on the bound so adjacent windows crossfade summing to 1. Every [DiT] Step log line appends gains=[..] when gating is active — check there first when a mix sounds one-sided.
  • Training side: Side-Step --timestep-window-min/max (rejection-resampled logit-normal; discrete mode filters the 8-step schedule) trains matching interval experts. T-LoRA: the high-noise expert overfits fastest — lower rank.
  • Inherits per-section constraints: runtime-only (no DoRA rescale), N× VRAM (Q4_0 knob), no basin re-base. runtime_lowrank is CLOBBERED to plain runtime (server + engine) — windows currently require the full-delta path; masked low-rank apply is untried future work.
  • VRAM trap (shipped bug, fixed same day): windows force runtime mode server-side, but getGlobalParams() gated adapterRuntimeQuant on the selected mode — a Merge-mode user with windows got full BF16 deltas (~8 GB/adapter on XL, 32 GB with 2 + reload churn) and their Merge-VRAM "Low" setting meant nothing. Quant now flows (and the Adapter VRAM knob shows) whenever windows are active. Any future knob that matters on a forced path must gate on the EFFECTIVE mode, not the selected one.
Show full SKILL.md (1,102 more words)Show less
Per-section hard constraints (violating any reproduces a shipped bug)
  1. Re-upload masks every step + after graph rebuilds — Golden rule 1.
  2. Mid-sampling alignment needs its own backend_sched_new; sharing dit->sched corrupts CUDA state → CUDA error: invalid argument (4e48176).
  3. Graph node budget AND scheduler hash-set must scale with adapter count: dit-graph.h:590 (graph_cap = 8192 + loras*4096) and dit.h:328-331 (sched_nodes = 8192 + adapters*4096, sized from the intended stack because m->loras isn't populated yet). Miss either → crash with N section adapters (e09d6a3).
  4. In HOTSTEP_SECTION_NOMASK debug mode, mask tensors must not be created at all — unconsumed graph inputs get no buffer, so uploading to them crashes (34dce60).

Key files

PathRole
engine/src/adapter-merge.hMerge mode: safetensors/PEFT/LyCORIS parsing, LoKr detection (:440), GPU merge graph (adapter_merge_on_backend :523, contains the DoRA weight-decompose for both GPU and host paths), delta compute (adapter_compute_delta :480), basin re-base (adapter_rebase_fetch :78), PEFT DoRA magnitude detection (lora_is_magnitude :183), per-module alpha_pattern parser (adapter_read_alpha_pattern :238 — REQUIRED for PEFT rank_pattern adapters, else per-module strength is alpha_global/rank_actual ≈ wildly wrong), entry adapter_merge (:1453)
engine/src/adapter-runtime.hRuntime mode: DiTLoRA* structs (:39-70), slot map dit_lora_slot (:105), stack delta-sum adapter_stage_delta (:173), quantized finalize (:804), adapter_load_runtime_stack (:1011), DoRA merge-only warnings (PEFT :299, LoKr :510)
engine/src/adapter-cancel.hg_adapter_cancel atomic + adapter_cancel_requested() — cooperative cancel of the ~17 s delta precompute; separate header so the server avoids ggml deps
engine/src/adapter-trt.hTensorRT IRefitter merge variant (#ifdef HOT_STEP_TRT)
engine/src/dit.hLoad orchestration: merge-vs-runtime, fusion skip (:420), once-per-stack re-base (:565), per-section per-adapter loads (:618), fail-don't-cache (:648). Fork hook: includes adapter-merge.h/adapter-runtime.h (:11/:13)
engine/src/dit-graph.hForward graph: dit_ggml_linear_lora (:45), DiTLoRASectionCtx (:80), per-adapter masks, NOMASK debug flag (:90)
engine/src/hot-step-sampler.hLive sampling loop: mask building/rebuilding, per-step re-uploads, P2 alignment
engine/src/hot-step-params.hg_hotstep_params sideband global (:205) + hotstep_adapter_stack_sig (:211); included by model-store.h (fork hook)
engine/src/model-store.{h,cpp}ModelKey adapter fields (model-store.h:90-102), hash/eq (model-store.cpp:48-94)
engine/src/pipeline-synth.cppBuilds the DiT cache key (:168-199)
engine/src/pipeline-synth-ops.cppHeader-anchored token→section map for P2 (:1419-1495); carries the hot-step-sampler.h fork hook (:9)
engine/tools/hot-step-server.cppThe REAL ace-server. ServerFields (:569), name→path resolution + g_hotstep_params fill (:1060-1162), cancel wiring (~:1214), job phase ADAPTER_PRECOMPUTE (:263)
server/src/services/generation/adapterSections.ts[Section]{k=v} directive parser + stripAdapterDirectives
server/src/services/generation/translateParams.tsUI params → aceReq adapter fields (:83-143)
server/src/routes/generate.tsLM-echo rebuild {...aceReq, ...lmFields} (:238, :320)
server/src/routes/adapters.tsFilesystem only: GET /api/adapters/browse, POST /api/adapters/scan (.safetensors listing)
ui/src/stores/globalParamsStore.ts, ui/src/components/global-bar/AdaptersDropdown.tsxAdapter stack, Sum/Blend, Adapter VRAM, Alignment Timing UI

Failure signatures

SymptomCause → fix
Adapters silently inaudible in per-section mode; mask logs show ~0Mask/input re-upload missing for steps 2..N or after a graph rebuild — Golden rule 1
Base-model output despite adapter selected, only after an earlier failed loadLoad failure cached as success under the adapter-bearing key — check dit.h:648 path (fixed 168dcb5)
Wrong adapter flavor after toggling merge↔runtime or changing quant/scale/group scalesMissing cache-key component — audit pipeline-synth.cpp:168-199 + model-store.cpp hash AND eq
CUDA error: invalid argument at the alignment stepAlignment shared the main scheduler — must be private (4e48176)
Crash/overflow when loading N section adaptersNode budget / sched hash-set not scaled with adapter count (dit-graph.h:590, dit.h:328)
Load stalls minutes at "quantize"Real Q4_K requested — use Q4_0 (Golden rule 6)
Crash reading delta tensor bytes in diagnosticsAssumed BF16; deltas may be Q4_0/Q8_0 (Golden rule 7)
Silence gaps between sections, later sections degenerateSomeone re-enabled self-attn isolation — REVERTED, Golden rule 4
Adapter VRAM / Alignment Timing knobs "do nothing"A ServerFields param dropped in an LM round-trip whitelist — Golden rule 8
Merge stack with re-base outputs only the LAST adapterRe-base applied per adapter instead of once per stack (dit.h:565)
Section adapters roughly twice as strong as expectedSection deltas loaded with stack scale instead of unit scale (dit.h:628)
DoRA adapter sounds wrong in runtime modeDoRA is merge-only; runtime can't express the multiplicative rescale — engine warns and applies plain LoRA (adapter-runtime.h:299 PEFT, :510 LoKr)
PEFT adapter with per-module ranks (rank_pattern in adapter_config.json) sounds weak/distortedalpha_pattern must be parsed per module (adapter_read_alpha_pattern) — the global lora_alpha over the actual per-tensor rank gives e.g. 128/512 instead of the trained 1024/512. Also requires adapter_config.json to sit NEXT TO the .safetensors (alpha fallback is rank ⇒ wrong strength without it)
All solvers/schedulers/guidance dead after an upstream sync (compiles fine)pipeline-synth-ops.cpp lost the hot-step-sampler.h include — run engine/verify-hooks.ps1

Institutional knowledge (from the departing lead engineer)

  • LoKR cross-base conversion — VALIDATED failure, VALIDATED fix. Converting adapters to non-turbo XL bases FAILS even though the bases are ~99% weight-identical. Root cause is basin-sensitivity (the adapter delta lands in a different loss basin), NOT weight drift. Fix: the β·(S−T) basin nudge (adapter-merge.h:78, applied in dit.h:578 for merge; adapter_runtime_rebase in adapter-runtime.h for runtime) — user-validated by ear in merge mode, 2026-07-14 ("works fantastically"); the runtime-mode port (7aac3c2, same day) is mathematically output-identical but awaits its own by-ear A/B. Never assume weight similarity implies adapter transferability.
  • Multi-adapter stacking (issue #72) — VALIDATED. Implemented as a sideband stack (g_hotstep_params.adapters) with read-from-pending merge accumulation (merge mode sources each tensor's current post-prior-merge value) and delta-sum in runtime mode (adapter_stage_delta — per-step cost and VRAM flat regardless of stack depth on the non-section path). CRITICAL: basin re-base once per stack, not per adapter (Golden rule 3).
  • Per-section masking — VALIDATED (shipped P1+P2). GGML clobbers input buffers, so masks/weights must be re-uploaded every step, never cached on device (Golden rule 1). Phase 1 = proportional token map; Phase 2 = alignment-based timing; the token→section map is now header-anchored (falls back to char-proportional when header count ≠ section count — the log line says which).
  • Regional self-attn isolation — VALIDATED failure. 0f3bf6d → reverted ee041e1. Do not re-attempt without a new design for cross-section coherence (Golden rule 4; safer untried idea in reference.md).
  • Cache-key gaps + load-failure caching — VALIDATED fix (168dcb5). Golden rule 2. The specific historical gaps were the time_embed/proj_in group scales and the runtime-mode marker; the failure-caching bug installed adapter-less models under adapter-bearing keys.
  • Cross-ARCH conversion (XL↔2B) — UNSOLVED; per-layer linear surgery is ruled out. Even with excellent per-layer alignment (CKA 0.90-0.99), independent per-layer linear transforms compound across 24 layers to ~5% end-to-end fidelity. Do not retry crop/pad, ridge, CCA, or Procrustes weight surgery. Distillation is the unblocked direction (details + open items in reference.md — several artifacts there are cited from local docs and unverified).

Debug tooling

  • HOTSTEP_BCAST_TEST=1 then run engine\build\Release\ace-synth.exe — numeric self-test of mask broadcast ops on the real backend, no models needed (engine/tools/ace-synth.cpp:160).
  • HOTSTEP_SECTION_NOMASK=1 — apply section deltas unmasked (dit-graph.h:90). Adapters audible ⇒ mask values are the problem; still silent ⇒ wiring/upload is the problem.
  • Logs: newest logs\<session>\ace_engine.log or logs\<session>\generations\gen_*.log. Useful greps: [Adapter-RT] Loaded, [Adapter-RT] P2:, header-anchored, [Adapter] Basin re-base, Per-section masking.
  • Audible validation is by ear only — ask the user to listen; do not use a browser agent for UI/audio verification. Do not delete experiment output files on your own prediction of their worth.

Deeper reading

  • reference.md (this folder) — merge/runtime internals, exact cache-key recipe, P2 alignment pipeline, cross-base/cross-arch research history, status ledger.
  • docs/plans/per-section-adapter-masking.md, docs/plans/cross-arch-adapter-conversion.md, docs/plans/multi-adapter-runtime-switching-handoff.md, docs/plans/2026-04-18-adapter-group-scales-investigation.md — local-only, gitignored; may be absent on a fresh clone. Their load-bearing content is distilled into this skill and reference.md.
  • engine/docs/ARCHITECTURE.md — engine-wide context (committed).

© scragnog, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in .claude/skills/adapter-system of scragnog/HOT-Step-CPP.

  • SKILL.md
  • reference.md

Open the folder on GitHubat commit 91e92a8

Compare with similar skills

Adapter System next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Adapter System compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Adapter System this skillscragnog/HOT-Step-CPP173—~6.3kAutomated safety check: PassMIT
Agent Sona Learning Optimizerruvnet/ruflo74k2 repos~516Automated safety check: PassMIT
Sentence-Transformers Training Routerhuggingface/skills11k1 repos~2.6kAutomated safety check: PassApache-2.0
Train RlOpenPipe/ART11k—~2.4kAutomated safety check: PassApache-2.0
Qwopus27b Rl TrainingR6410418/Jackrong-llm-finetuning-guide1.7k—~830Automated safety check: PassApache-2.0
Dataset Evaluationawslabs/agent-plugins9151 repos~1.3kAutomated safety check: PassApache-2.0

Similar skills

  • Agent skill for sona-learning-optimizer - invoke with $agent-sona-learning-optimizer

    74k GitHub starsUsed in 2 repos~516 tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Routes a sentence-transformers training task to the right model type and required reference docs and example scripts, covering bi-encoders, rerankers, sparse and multi-vector models.

    11k GitHub starsUsed in 1 repo~2.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Train Rl

    OpenPipe/ART

    RL training reference for the ART framework. An agent skill from OpenPipe/ART.

    11k GitHub stars~2.4k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Qwopus27b Rl Training

    R6410418/Jackrong-llm-finetuning-guide

    Prepare, validate, launch-plan, monitor, resume, and stop configurable Qwopus 27B reinforcement-learning workflows for GRPO or GSPO.

    1.7k GitHub stars~830 tokensUpdated 3 mo ago
    AI & LLM EngineeringAuto-check passed
  • Dataset Evaluation

    awslabs/agent-plugins

    Official

    Validates dataset formatting and quality for SageMaker model fine-tuning (SFT, DPO, or RLVR).

    915 GitHub starsUsed in 1 repo~1.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Train Sft

    OpenPipe/ART

    SFT training reference for the ART framework. An agent skill from OpenPipe/ART.

    11k GitHub stars~2.9k tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from scragnog/HOT-Step-CPP

All 18 skills in this repo
  • Ear Test Scoresheet

    scragnog/HOT-Step-CPP

    The standard way to run a listening test in HOT-Step - a local HTML score sheet next to the renders where Rob plays each track, scores it 1-5 on named criteria, and the page charts the two score…

    173 GitHub stars~1.9k tokensUpdated 2 days ago
    Auto-check passed
  • Engine Performance

    scragnog/HOT-Step-CPP

    Explains where HOT-Step generation time goes (LM/DiT/VAE), how the TensorRT paths activate, how to benchmark from logs, and which knobs trade quality for speed.

    173 GitHub stars~4.9k tokensUpdated 2 days ago
    Auto-check passed
  • Mm3 Backend

    scragnog/HOT-Step-CPP

    Maps HOT-Step's native MiniMax-Music3 backend — engine port modules, endpoints, server/UI integration, parity/fixture infrastructure, and the hard-won trap list.

    173 GitHub stars~4.5k tokensUpdated 2 days ago
    Auto-check passed
  • Mm3 Lm Adapter Training

    scragnog/HOT-Step-CPP

    The validated recipe for training MiniMax-Music3 planner-LM style adapters (artist/album clones) with ace-train mm3-lm-train and the Training Studio.

    173 GitHub stars~4k tokensUpdated 2 days ago
    Auto-check passed
  • Release Process

    scragnog/HOT-Step-CPP

    Runbook for cutting and publishing a HOT-Step CPP release via a v git tag that triggers the multi-platform CI build and drafts a GitHub Release.

    173 GitHub stars~5.1k tokensUpdated 2 days ago
    Auto-check passed
  • Upstream Sync

    scragnog/HOT-Step-CPP

    Safely pulls upstream acestep.cpp changes into the HOT-Step engine fork without destroying its integration hooks.

    173 GitHub stars~5k tokensUpdated 2 days ago
    Auto-check passed

Works with

Questions about Adapter System

What does Adapter System do?

Explains how HOT-Step's LoRA/LoKr adapter system loads, merges, caches, stacks, and regionally masks adapters at runtime, including hard-won failure modes. Adapter System is an agent skill from scragnog/HOT-Step-CPP. Explains how HOT-Step's LoRA/LoKr adapter system loads, merges, caches, stacks, and regionally masks adapters at runtime, including hard-won failure modes.

When should I use Adapter System?

Adapter System fits situations like: working on adapter loading/merging; multi-adapter stacking; per-section adapter masking; adapter cache keys.

How do I install Adapter System in Claude Code?

Run `npx skills add scragnog/HOT-Step-CPP --skill adapter-system -a claude-code`. Or copy the skill folder (.claude/skills/adapter-system in scragnog/HOT-Step-CPP) into .claude/skills/adapter-system in your project. Claude Code loads it when a task matches its description.

How do I install Adapter System in Codex?

Run `npx skills add scragnog/HOT-Step-CPP --skill adapter-system -a codex`. Or copy the skill folder (.claude/skills/adapter-system in scragnog/HOT-Step-CPP) into .agents/skills/adapter-system in your project. Codex loads it when a task matches its description.

Can I use Adapter System in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add scragnog/HOT-Step-CPP --skill adapter-system -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/adapter-system, .gemini/skills/adapter-system, .github/skills/adapter-system and .opencode/skills/adapter-system in your project.

What does Adapter System need to run?

Going by SKILL.md and its folder, Adapter System needs the command-line tools its instructions call (cmake, npx, npm and git). Our summary lists: Node.js.

Does Adapter System access the network?

SKILL.md contains no URLs. Its commands use npx, npm and git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Adapter System safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Adapter System use?

Adapter System is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Adapter System use?

About 6.3k tokens (SKILL.md is roughly 25k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Adapter System?

Skills that share tags, products or a category with Adapter System: Agent Sona Learning Optimizer (ruvnet/ruflo, 74k stars), Sentence-Transformers Training Router (huggingface/skills, 11k stars), Train Rl (OpenPipe/ART, 11k stars) and Qwopus27b Rl Training (R6410418/Jackrong-llm-finetuning-guide, 1.7k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Adapter System?

scragnog (a GitHub user) maintains it in scragnog/HOT-Step-CPP, which has 173 GitHub stars. The repository holds 18 skills in this directory. The repository was last updated on October 7, 2026.

Source: scragnog/HOT-Step-CPP on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.