Agent skill

Model Management

by scragnog in scragnog/HOT-Step-CPP

Explains how model files, checkpoints, GGUF quantization, and the Model Manager work in HOT-Step CPP.

MITAuto-check passedAI & LLM Engineering

Install Model Management

skills CLI
$ npx skills add scragnog/HOT-Step-CPP --skill model-management -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install scragnog/HOT-Step-CPP model-management --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/scragnog/HOT-Step-CPP.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/model-management .claude/skills/model-management && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
model-management
GitHub stars
173
Token cost
~6.6k tokens
SKILL.md length
2,691 words
Files
2
Skills in repo
18
Repo updated
First seen
Licence
MIT

At a glance

Explains how model files, checkpoints, GGUF quantization, and the Model Manager work in HOT-Step CPP.

  • Works in 8 steps: Two independent registries exist — don't… → GGUF files must sit in the models root,… → Never hand-edit or hand-build a GGUF's… → …
  • Adding/converting/quantizing models
  • SKILL.md covers Terminology (read first), When to use this skill, Golden rules and Directory layout (expected), plus 11 more sections
  • Calls node, python and cmake; reaches huggingface.co

What it does

Model Management is an agent skill from scragnog/HOT-Step-CPP. Explains how model files, checkpoints, GGUF quantization, and the Model Manager work in HOT-Step CPP. Use when adding/converting/quantizing models, debugging "no GGUF models found" or missing-model/wrong-model failures, working on the model download service or Model Manager UI, or answering which component (LM/DiT/VAE/text-encoder) needs which file.

Its SKILL.md is about 6.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `reference.md`).

It sits in AI & LLM Engineering, covering LLM inference and serving. It works with llama.cpp, C++ and Qwen. The repository describes itself as: Turn dials. Summon bangers! NOW WITH MORE C++! Local AI music generation powered by GGML. The licence is MIT.

When your agent uses it

  • Adding/converting/quantizing models
  • Debugging no GGUF models found
  • Missing-model/wrong-model failures
  • Working on the model download service

Example prompts

  • “no GGUF models found”
  • “Use the model-management skill to explain how model files, checkpoints, GGUF quantization, and the Model Manager work in HOT-Step CPP”
  • “/model-management”

Requirements

  • Python 3

Workflow steps

8 steps, taken from the first numbered list in SKILL.md.

  1. Two independent registries exist — don't confuse them. The Node-side curated catalogue (server/src/data/model-registry.json) controls what…
  2. GGUF files must sit in the models root, not a subfolder. The engine scans only root-level .gguf files…
  3. Never hand-edit or hand-build a GGUF's tensor set for a recognized architecture. A recognized-arch GGUF with a missing tensor kills the…
  4. Quantize from BF16 sources only, and never quantize the VAE. quantize.exe reads BF16 GGUF input; VAE-arch tensors and small/critical…
  5. Rebuild rules apply here too: any change to engine/src/model-registry.h, model-store.h, gguf-weights.h, etc. means rebuild via…
  6. Runtime DLLs do not go in models/. Catalogue entries with role: "runtime" (cuBLAS, cudart, ONNX Runtime, cuDNN) install next to…
  7. A model that only exists on this machine does not exist. Adding a model the app resolves at runtime is a three-part act: upload the…
  8. Model/adapters directory paths are restart-required config. ACESTEPCPP_MODELS / ACESTEPCPP_ADAPTERS env vars override defaults /models and…

What it can do on your machine

Read from SKILL.md and the folder at commit 91e92a8. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • node
    • python
    • cmake
    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • huggingface.co

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Model Management loads about 6.6k tokens when it runs. Until then it costs about 92 tokens; SKILL.md has 2,691 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~92
When it runs · the whole SKILL.md, loaded when a task matches
~6.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from scragnog/HOT-Step-CPP at commit 91e92a8, republished under its MIT licence (© scragnog). 2,691 words, ~6,572 tokens.

Download SKILL.mdSave it as .claude/skills/model-management/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
model-management
description
Explains how model files, checkpoints, GGUF quantization, and the Model Manager work in HOT-Step CPP. Use when adding/converting/quantizing models, debugging "no GGUF models found" or missing-model/wrong-model failures, working on the model download service or Model Manager UI, or answering which component (LM/DiT/VAE/text-encoder) needs which file.

Model Management — Models, Checkpoints & Quantization

All paths are repo-relative to the repo root (d:\Ace-Step-Latest\hot-step-cpp). All commands are Windows PowerShell (use ; to chain, never &&).

Terminology (read first)

  • LM — the language model (Qwen3-based) that turns caption+lyrics text into audio codes. Files: acestep-5Hz-lm-{0.6B,1.7B,4B}-*.gguf.
  • DiT — Diffusion Transformer, the denoising model that generates audio latents. Files: acestep-v15-*.gguf. One DiT GGUF is self-contained (also carries the condition encoder, FSQ tokenizer/detokenizer, silence_latent, null_condition_emb).
  • Text encoder — Qwen3 embedding model that encodes the caption for the DiT. Files: Qwen3-Embedding-0.6B-*.gguf.
  • VAE — decodes latents to 48 kHz audio (and encodes audio to latents for cover/repaint/extend). Files: vae*.gguf / scragvae*.gguf.
  • PP-VAE — optional post-processing VAE ("polish" re-encode pass). Files: pp-vae-*.gguf.
  • GGUF — the GGML binary weight format the C++ engine loads (mmap'd). Each GGUF declares its role in the header key general.architecture.
  • Quant — reduced-precision weight encoding (Q4_K_M, Q8_0, MXFP4, ...) to shrink VRAM/disk. Produced from BF16 GGUFs by engine/tools/quantize.cpp.
  • Adapter — LoRA/LoKr fine-tune delta (.safetensors), lives in adapters/, not models/.
  • Model Manager — the "Get More Models" modal in the UI + Node download service that fetches curated files from HuggingFace.

When to use this skill

  • Installing, converting, or quantizing model files; deciding where a file must live.
  • Debugging: engine won't start, /synth unavailable, model missing from dropdowns, crash-loops, corrupt downloads.
  • Modifying the Model Manager (route server/src/routes/modelManager.ts, service server/src/services/modelDownloadService.ts, UI ui/src/components/model-manager/), or the catalogue server/src/data/model-registry.json.

Golden rules

  1. Two independent registries exist — don't confuse them. The Node-side curated catalogue (server/src/data/model-registry.json) controls what is downloadable; the C++ engine's startup scan (engine/src/model-registry.h) controls what is usable. A file can show "installed" in the Model Manager yet be invisible to generation dropdowns (and vice versa). WHY: they use different detection logic (filename presence vs GGUF-header architecture) and different directory depths.
  2. GGUF files must sit in the models root, not a subfolder. The engine scans only root-level .gguf files (engine/src/model-registry.h:234-267), but the Node installed-check also scans one subdir level (modelDownloadService.ts:177-208). A GGUF in a subfolder = "installed" in the UI, dead to the engine. WHY: silent classic confusion — no error anywhere.
  3. Never hand-edit or hand-build a GGUF's tensor set for a recognized architecture. A recognized-arch GGUF with a missing tensor kills the whole ace-server process: gf_load_tensor() prints [GGUF] FATAL: tensor 'x' not found and calls exit(1) (engine/src/gguf-weights.h:164-168). Node respawns it, so a bad model file mid-request looks like a random engine crash-loop.
  4. Quantize from BF16 sources only, and never quantize the VAE. quantize.exe reads BF16 GGUF input; VAE-arch tensors and small/critical tensors (silence_latent, scale_shift_table, null_condition_emb, 1-D tensors, text-enc embed_tokens) are deliberately never quantized (engine/tools/quantize.cpp:89-109). WHY: quantizing these destroys audio quality or breaks generation outright.
  5. Rebuild rules apply here too: any change to engine/src/model-registry.h, model-store.h, gguf-weights.h, etc. means rebuild via dev-rebuild.bat at repo root — never engine/build.cmd directly (you cannot reliably tell whether the app is running; Node auto-respawns ace-server, and killing it uncleanly causes an infinite respawn + file-lock loop). Never cmake --clean-first (20+ min CUDA recompile).
  6. Runtime DLLs do not go in models/. Catalogue entries with role: "runtime" (cuBLAS, cudart, ONNX Runtime, cuDNN) install next to ace-server.exe (modelDownloadService.ts:110-118). WHY: missing DLLs there are the #1 cause of the "crashed 3 times within 30s — giving up" loop.
  7. A model that only exists on this machine does not exist. Adding a model the app resolves at runtime is a three-part act: upload the weights to Hugging Face, add the entry to server/src/data/model-registry.json, and put it in a pack if a feature requires it. Do all three or none — a feature gated on a file that was never uploaded looks fine here and is dead for every user. Verify with node server/scripts/check-release-prereqs.mjs. WHY: MM3 training shipped in v1.3 demanding mm3-rvq-*.gguf and mm3-enc-*.gguf that had never left the dev box (#137, fixed 96d442fb).
  8. Model/adapters directory paths are restart-required config. ACESTEPCPP_MODELS / ACESTEPCPP_ADAPTERS env vars override defaults <repo>/models and <repo>/adapters (server/src/config.ts:53-54,100-101); the engine gets them as spawn-time --models/--adapters flags.

Directory layout (expected)

models/                                  # config.aceServer.models (default <repo>/models)
  acestep-v15-*.gguf                     # DiT (arch "acestep-dit"): base/sft/turbo/merge x BF16/Q8_0/Q6_K/Q5_K_M/Q4_K_M/MXFP4/NVFP4...
  acestep-5Hz-lm-{0.6B,1.7B,4B}-*.gguf   # LM (arch "acestep-lm")
  Qwen3-Embedding-0.6B-*.gguf            # Text encoder (arch "acestep-text-enc")
  vae-*.gguf, scragvae-*.gguf            # VAE (arch "acestep-vae"); ScragVAE/Regrind = drop-in decoder variants
  pp-vae-*.gguf                          # PP-VAE (arch "pp-vae")
  vae-*.safetensors                      # safetensors VAE, classified by filename prefix only
  vae-*.onnx                             # ONNX VAE — DECODER-ONLY (see failure table)
  <name>/                               # HF safetensors checkpoint dir: config.json + model.safetensors
                                        #   (or model.safetensors.index.json for sharded) — classified by config.json content
  onnx/                                 # ONNX Runtime / TensorRT model dirs (config.aceServer.onnxDir)
  supersep/*.onnx                       # stem-separation nets (Cover/Stem Studio)
  whisper/ggml-*.bin                    # whisper.cpp models (config.ts:253-255)
adapters/                                # config.aceServer.adapters
  <name>.safetensors                    # ComfyUI single-file LoRA (alpha baked in)
  <name>/adapter_model.safetensors      # PEFT directory format

Which component needs which model

Verified in engine/src/model-store.h:68-81 (ModelKind comments):

ComponentModel fileNotes
LM (MODEL_LM)acestep-5Hz-lm-*.ggufONE shared instance for generate + ace-understand — enforced by identical ModelKey (model-store.h:17-22)
Text encoder (MODEL_TEXT_ENC)Qwen3-Embedding-*.gguf
Cond-enc + DiT + FSQ tok/detokthe same acestep-v15-*.ggufSelf-contained; also holds silence_latent + null_condition_emb (model-store.h:112-118)
VAE encode + decodevae*.gguf (has encoder.* and decoder.*)
PP-VAE polishpp-vae-*.ggufRequest flag pp_vae_reencode (engine/src/request.h:146); availability = GET /api/models/pp-vae scans for pp-vae*.gguf (server/src/routes/models.ts:50-65)
ORT/TRT accelerationmodels/onnx/ subdirsMODEL_*_ORT kinds
SuperSep stemsmodels/supersep/*.onnx+ ONNX Runtime/cuDNN DLLs beside ace-server.exe
Whisper transcriptionmodels/whisper/ggml-*.bin+ tools/whisper/whisper-cli.exe (config.ts:254-255)

Synth pipeline needs DiT + Text-Enc + VAE simultaneously. Missing any one → /synth unavailable warning (server stays up LM-only if an LM exists), or exit 1 if no LM either (engine/tools/hot-step-server.cpp:2600-2616 — the compiled server; the same block exists in the UNCOMPILED reference copy ace-server.cpp:1686-1711). Request-level selection: synth_model, lm_model, vae are filenames resolved against the engine's scanned registry; empty string = first matching entry (engine/src/request.h:125-142).

Procedure: quantize a BF16 GGUF

powershell
# From repo root. Binary lives at engine\build\Release\quantize.exe
.\engine\build\Release\quantize.exe <input-BF16.gguf> <output.gguf> <TYPE>
# Example:
.\engine\build\Release\quantize.exe models\acestep-v15-turbo-BF16.gguf models\acestep-v15-turbo-Q4_K_M.gguf Q4_K_M
  • Valid TYPEs (case-insensitive, quantize.cpp:8): Q2_K Q3_K_S Q3_K_M Q3_K_L Q4_K_S Q4_K_M Q5_K_S Q5_K_M Q6_K Q8_0 NVFP4 MXFP4. IQ3/IQ4 quants seen on disk are not producible by this tool.
  • Mixed-precision policy mirrors llama-quantize: "important" tensors (v_proj, down_proj; L variants add o_proj) bumped one tier; embed_tokens always Q6_K (Q8_0 for Q8_0/NVFP4/MXFP4) (quantize.cpp:40-54,74-86).
  • Streaming write, low memory. Prints Quantized N/M tensors + compression ratio.
  • Output goes straight into models\ root → picked up on next engine restart.

Procedure: convert HF safetensors → BF16 GGUF (engine/convert.py)

  • No CLI args. Hardcoded: reads checkpoint dirs from engine/checkpoints/, writes GGUFs to engine/models/ (convert.py:14-16). Neither directory exists in this working tree — create engine\checkpoints\, put the HF checkpoint dir inside, run it, then move the output GGUF to repo-root models\.
  • Classification is by checkpoint directory name: acestep-5Hz-lm* → LM, acestep-v15* → DiT, Qwen3-Embedding* → text-enc, and exactly vae → VAE (convert.py:55-64). Skips outputs that already exist.
  • Alternative: the engine loads safetensors checkpoint dirs directly (drop <name>/ with config.json + model.safetensors into models/) — conversion is optional. Sharded (model.safetensors.index.json) and diffusers (diffusion_pytorch_model.safetensors) layouts supported (engine/src/weight-source.h, engine/src/model-registry.h:328-375).

Procedure: convert ComfyUI int8 DiT safetensors → Q8_0 GGUF (engine/convert-comfy-int8.py)

For ComfyUI comfy_quant int8 DiT checkpoints (int8 .weight + F32 .weight_scale scalar or per-row + .comfy_quant JSON tensor per layer), including ConvRot files ("convrot": true + convrot_groupsize). Per-tensor and per-row int8 grids are exactly representable in Q8_0 (block scale = tensor/row scale), so weights are repacked bit-faithfully — no dequant/requant round trip. Needs a donor GGUF of the same architecture (any convert.py-produced acestep-v15-*.gguf) to supply silence_latent and the acestep.* config KVs, which ComfyUI files lack. Aborts on any tensor-shape mismatch vs the donor.

powershell
python engine\convert-comfy-int8.py <comfy.safetensors> models\<matching-donor>-BF16.gguf models\<out>-Q8_0.gguf --name <general.name>

ConvRot handling: rotated decoder.* weights stay rotated and are recorded in GGUF KV acestep.convrot_map (name:group;...); the engine applies the matching group-wise Hadamard rotation to that linear's activations at inference (dit.h load + dit-graph.h/dit-alignment-graph.h, commit 182faef). Rotated encoder/tokenizer/detokenizer weights are dequantized + unrotated to BF16 offline (run once per generation — not worth graph wiring). --no-runtime-rotation builds an all-BF16 unrotated reference GGUF of the same quantized model, used for same-seed A/B validation of the engine rotation path. Adapter merge mode is refused on ConvRot models (deltas are unrotated); runtime adapter mode works (deltas consume raw activations).

ConvRot cardinal rule — any code reading ConvRot base weights for unrotated-space math must unrotate them first (convrot_transform_rows in engine/src/convrot.h, fast radix-4, self-inverse). Violation signature: generation "succeeds" but output is garbled full-band noise. First instance: the runtime basin re-base nudged deltas with β·(S−T) using rotated T — fixed in adapter_runtime_rebase (commit 22820ae) by unrotating T for acestep.convrot_map tensors. Audit any future weight-reader (TRT export, distills, external merge scripts) against this.

Producing ConvRot files from a local checkpoint: pip install convert_to_quant (needs torch+CUDA, triton-windows) then ctq -i <model.safetensors> -o <out.safetensors> --int8 --scaling_mode row --convrot --dynamic_convrot --comfy_quant --save-quant-metadata (~35 min for a 5B XL on an RTX 5090, learned rounding included).

First applied 2026-07-15: hrktxz xl_sft_turbo (plain int8) → acestep-v15-xl-sft-turbo-comfy-int8-Q8_0.gguf; merge-base-sft-turbo-xl-thirds self-quantized with real ConvRot → ...-convrot-Q8_0.gguf (+ ...-convrot-ref-BF16.gguf reference). Numerical parity: rotated-path error 0.9% vs original F32 weights; skipping rotation → ~140% (i.e. rotation is load-bearing).

Procedure: add a model manually

  1. Copy the .gguf into models\ root (not a subfolder — golden rule 2).
  2. Restart the engine (dev-rebuild.bat restarts everything, or restart the app). The scan runs only at ace-server startup.
  3. Check the newest logs\<session>\ace_engine.log for [Registry] <file> -> DiT (or LM/VAE/...). A WARNING: skipping X (unknown architecture) means the GGUF header lacks a recognized general.architecture (acestep-lm|acestep-dit|acestep-text-enc|acestep-vae|pp-vae, model-registry.h:99-130).
  4. The file now appears in GET /api/models (Node proxies engine /props; buckets lm, embedding, dit, vae — hot-step-server.cpp:2326-2329 — the compiled server, not the uncompiled ace-server.cpp).

Procedure: publish a new model so users can get it

Local conversion/quantization is only half the job. Until these steps are done the model does not exist for anyone but you.

  1. Upload the weights. huggingface_hub is installed; the token lives in ~/.cache/huggingface/token (account scragnog). Ask the user before pushing to a public repo — it is outward-facing and hard to walk back.

    python
    from huggingface_hub import HfApi
    HfApi().upload_file(path_or_fileobj='models/mm3/<file>.gguf', path_in_repo='<file>.gguf',
                        repo_id='scragnog/<repo>', repo_type='model',
                        commit_message='Add <file>')
  2. Add the registry entry to server/src/data/model-registry.json — id, filename, role, subdir, displayName, quant, exact sizeBytes (the downloader validates size ±5%), repo, description, tags, and a companions LICENSE entry if the weights carry one. The JSON round-trips exactly under json.dumps(indent=2, ensure_ascii=False), so it can be edited programmatically without reformatting the whole file.

  3. Add it to a pack if a feature needs it, and check the reverse: a feature that resolves files by prefix scan rather than by registry id (e.g. resolveMm3TrainModels takes the newest mm3-rvq-*.gguf on disk) will not be satisfied unless a published filename matches the prefix.

  4. Credit the author in the HF model card if the weights are not ours, and keep the upstream licence. Community encoders and adapters are other people's work.

  5. Verify: node server/scripts/check-release-prereqs.mjs — checks every entry resolves on HF at the claimed size, packs reference real ids, and runtime data files are packaged. Exit 1 = do not ship.

Procedure: drive the Model Manager via API

Routes in server/src/routes/modelManager.ts, mounted at /api/model-manager:

powershell
Invoke-RestMethod http://localhost:3000/api/model-manager/registry            # catalogue + installed flags (dev app)
Invoke-RestMethod -Method Post -Uri http://localhost:3000/api/model-manager/download -ContentType 'application/json' -Body '{"fileId":"<id>"}'
# GET /downloads = SSE progress stream; POST /download/<jobId>/cancel | /resume; DELETE /files/<filename>

Download mechanics: HuggingFace URL https://huggingface.co/{repo}/resolve/main/{repoPath || filename}, resume via HTTP Range + .part file, 3 attempts (0/2s/5s), validation before rename (size ±5%, MZ header for .dll, GGUF magic for .gguf) — modelDownloadService.ts:352-473. Details and data shapes: reference.md.

Show full SKILL.md (1,065 more words)Show less

Key files

PathRole
engine/src/model-registry.hEngine startup scan/classification of --models and --adapters dirs
engine/src/model-store.hRefcounted VRAM ownership; EVICT_STRICT (default) vs EVICT_NEVER (--keep-loaded); ModelKey caching incl. adapter extras
engine/src/gguf-weights.hmmap GGUF loader; truncation guard; FATAL exit on missing tensor
engine/src/safetensors.h, engine/src/weight-source.hsafetensors parser + format-agnostic layer (GGUF/safetensors)
engine/tools/quantize.cpp → engine/build/Release/quantize.exeBF16 GGUF → K-quant/FP4 GGUF
engine/convert.pyHF safetensors checkpoint dir → BF16 GGUF (hardcoded dirs)
engine/tools/ace-server.cppStartup validation, /props endpoint
server/src/config.tsaceServer.models/adapters/onnxDir, keepLoaded, warm-on-startup, whisper paths
server/src/services/modelDownloadService.tsDownload jobs, resume, validation, installed-check, variant filtering
server/src/routes/modelManager.ts/api/model-manager/* REST + SSE
server/src/routes/models.ts/api/models (proxies engine /props), /api/models/pp-vae
server/src/data/model-registry.jsonCurated catalogue: 152 files, 9 packs
server/src/index.tsace-server spawn/respawn limiter (152-156, 284-308); first-launch CUDA DLL bootstrap (318-380)
ui/src/components/model-manager/Modal UI: ModelManagerModal.tsx, ModelCatalogueTab.tsx (7 tabs), ModelRow.tsx, StarterPackCard.tsx, DownloadProgressBar.tsx, useModelRegistry.ts, useDownloadStream.ts

Failure signatures

SymptomCauseFix
[Server] ERROR: no models found + engine exit 1; Node retries 3x then gives upEmpty/wrong models dir (ACESTEPCPP_MODELS), or nothing classifiable — and no MM3 weights either. Since the issue-#118 fix, MM3 weights (mm3-*.gguf at root or in mm3/) keep the server alive MM3-only ("No ACE-Step models … continuing MM3-only")Point at the right dir / install models; restart
[Registry] WARNING: skipping X (unknown architecture)GGUF header lacks a recognized general.architectureConvert via convert.py, or it's not an ACE-Step GGUF
[Server] WARNING: /synth unavailable, missing: VAE (etc.)Partial install — synth needs DiT+Text-Enc+VAE togetherDownload the missing role (Model Manager quick-start pack)
[GGUF] FATAL: '<f>' is truncated or corrupt ... file is only N bytesInterrupted download / prematurely renamed .partDelete and re-download
[GGUF] FATAL: tensor 'x' not found then process deathRecognized arch, wrong/incomplete tensor set — kills ace-server mid-requestRemove the bad GGUF
Download "completes" then Invalid GGUF header — got "<!DO"HuggingFace served an HTML error page (auth/rate-limit/404)Retry; check the repo/path in the catalogue entry
Size mismatch: expected X MB, got Y MBCatalogue sizeBytes drift vs repo file, or corrupt transferRe-download; fix sizeBytes in model-registry.json if repo file changed
Crash-loop "3 times within 30s" + missing-DLL hintcuBLAS/cudart DLLs absent beside ace-server.exe (CUDA variant)Model Manager "CUDA Runtime" pack; first-launch bootstrap normally handles it
Model shows "installed" in Model Manager but absent from generation dropdownsGGUF in a subdir (Node scans subdirs, engine scans root only), unknown arch, or engine down (aceServerDown: true)Move to models root / check engine log
Cover/repaint fails while text2music works, ONNX VAE selectedONNX VAEs are decoder-only; VAE encode requires a non-ONNX VAE (registry_find_non_onnx, model-registry.h:62-85)Install a GGUF/safetensors VAE alongside
Download of the DreamVAE entry fails instantlyThat catalogue entry has no repo field → URL contains undefinedTreat as local/display-only (see below). Regrind entries were fixed 2026-07-29 (now download from mdmachine/ACEStep-XL-Regrind-V1)

Institutional knowledge

  • VALIDATED (production incident, 2026-08-31): the registry and Hugging Face are the distribution boundary, not models/. MM3 training in v1.3 required mm3-rvq-*.gguf + mm3-enc-*.gguf, which existed only on the dev machine and in neither the catalogue nor any HF repo, so the Training Studio asked every user for files that could not be obtained (#137). Fixed 96d442fb by publishing both and adding a minimax-music3-training pack (which also carries mm3-depth-f16 — the trainer needs f16 depth while every generation pack ships a quantised one). node server/scripts/check-release-prereqs.mjs now gates it.

  • The RVQ encoder is not ours. mm3-rvq-53kpooled-f32.gguf is PurpleOrc's open-rvq (SimpleTuner v4 architecture, 53k-track corpus), mirrored to our repo under the same MiniMax-Music3 community terms with credit in the model card. Codes are encoder-specific: an adapter trained on these codes must keep using this encoder, so replacing it means re-exporting every code cache.

  • VALIDATED — subdir-GGUF blind spot: engine scans root-only for .gguf; Node installed-check scans one subdir level. This mismatch is a real, recurring "installed but not selectable" source (code cited in golden rule 2).

  • VALIDATED — truncation guard exists for a reason: the byte-range check in gf_load() (gguf-weights.h:132-150) was added specifically because truncated downloads used to segfault deep in cuMemcpyHtoDAsync. Keep it if touching the loader.

  • CORRECTED 2026-08-31 — the repo-less entry is fixed. vae-dreamvae-onnx used to have no repo field, so startDownload built https://huggingface.co/undefined/... and failed; it now carries repo: daydreamlive/DreamVAE + repoPath: onnx/model.onnx and resolves (verified by check-release-prereqs.mjs). The failure mode is still worth knowing: an entry missing repo is display-only, visible in the catalogue and undownloadable, which is exactly the class of fault the prereq check now catches. (The Regrind V9b entries had the same gap until 2026-07-29 — all 8 Regrind VAEs now carry repo: mdmachine/ACEStep-XL-Regrind-V1 + repoPath: vae/...; that repo keeps files in vae//dit//lora/ subfolders, so repoPath is mandatory for anything added from it.)

  • VALIDATED — installed-check is name-only: filename presence, no size/hash (modelDownloadService.ts:177-208). A stale/partial file with the right name shows "installed".

  • VALIDATED — --keep-loaded trade-off: default EVICT_STRICT reloads models per request; ACESTEPCPP_KEEP_LOADED=1 → --keep-loaded (EVICT_NEVER) avoids ~17 s LoKr adapter precompute per request but pins ~13 GB VRAM (config.ts:124-132, index.ts:177-184). Warm-on-startup env vars (ACESTEPCPP_WARM_DIT/_VAE/_ADAPTER) only fire when keepLoaded is on.

  • VALIDATED — DiT instances cache per adapter combo: ModelKey includes adapter path/scale, per-group scales, basin re-base (rebase_source/rebase_beta), and multi-adapter adapter_stack signature (model-store.h:83-103) — different values = distinct cached DiTs, each costing VRAM under keep-loaded.

  • VALIDATED — /api/models fallback shape is not a faithful mirror: when the engine is down, Node returns buckets { dit, lm, vae, understand } + aceServerDown: true (models.ts:28-36), but the live engine /props sends { lm, embedding, dit, vae } (ace-server.cpp:1512-1515) — no understand bucket exists, and the fallback lacks embedding.

  • VALIDATED — cleanupJobs() doc-drift: its comment says "older than 60s" but it deletes all terminal jobs immediately, and no route calls it (modelDownloadService.ts:341-348).

  • VALIDATED — speculative-decoding draft LM is DISABLED: ACESTEPCPP_DRAFT_LM plumbing remains (config.ts:117-122) but GGML per-call overhead negated the speedup; auto-detect commented out.

  • VALIDATED — "convrot" HF uploads may be mislabeled plain int8: all files in hrktxz/ACE_Step_1.5_ComfyUI_int8_convrot (checked 2026-07-15) contain only {"format": "int8_tensorwise"} per-tensor-scale layers — no Hadamard rotation metadata anywhere, despite the repo name. Real ConvRot (QuaRot-style group-wise Hadamard, ComfyUI ≥0.27) would need "convrot": true + convrot_groupsize in the .comfy_quant JSON and runtime activation rotation. Our vendored GGML already ships FWHT kernels (CUDA/Vulkan/CPU) + GGML_HINT_SRC0_IS_HADAMARD (ggml.h:444, unused by engine code) if we ever want native support. Check the safetensors header before believing a quant-format claim.

  • UNVALIDATED: provenance of on-disk IQ3/IQ4 quant files (not producible by quantize.cpp — presumably made with external tooling); whether neural-codec/mp3-codec binaries take model paths (not inspected).

Deeper reading

  • reference.md (this folder) — engine scan order/classification details, ModelKey/VRAM policy, download-service internals, catalogue statistics, API data shapes, quantize policy table.
  • engine/docs/ARCHITECTURE.md — engine internals, request JSON, generation modes (committed).
  • README.md — install/build; states safetensors DiT/LM/Text-Enc/VAE all loadable and BF16 safetensors bit-perfect vs BF16 GGUF.
  • docs/plans/2026-05-03-model-manager-design.md — original Model Manager design (gitignored, local-only, may be absent; code has drifted from it — trust code).

© scragnog, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in .claude/skills/model-management of scragnog/HOT-Step-CPP.

  • SKILL.md
  • reference.md

Open the folder on GitHubat commit 91e92a8

Compare with similar skills

Model Management next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Model Management compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Model Management this skillscragnog/HOT-Step-CPP173—~6.6kAutomated safety check: PassMIT
Qwen Mtp GgufR6410418/Jackrong-llm-finetuning-guide1.7k—~1.7kAutomated safety check: PassMIT
Add Modelguoqingbao/xinfer334—~4.2kAutomated safety check: NotesMIT
Resolvealexziskind1/model-shelf130—~792Automated safety check: PassMIT
Test Modelguoqingbao/xinfer334—~2.6kAutomated safety check: PassMIT
Add New ModelJakeATX/llamAmpere149—~4.1kAutomated safety check: PassMIT

Similar skills

  • Qwen Mtp Gguf

    R6410418/Jackrong-llm-finetuning-guide

    Complete agent-ready workflow for Qwen-family MTP or nextn GGUF conversion and release.

    1.7k GitHub stars~1.7k tokensUpdated 3 mo ago
    AI & LLM EngineeringAuto-check passed
  • Add Model

    guoqingbao/xinfer

    Adapt and port new LLM model architectures to this xinfer project.

    334 GitHub stars~4.2k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check: notes
  • Resolve

    alexziskind1/model-shelf

    Always resolve Hugging Face models via model-shelf before any download.

    130 GitHub stars~792 tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Test Model

    guoqingbao/xinfer

    Test LLM models served by xinfer for correctness, output quality, and performance.

    334 GitHub stars~2.6k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Add New Model

    JakeATX/llamAmpere

    Guided workflow for adding a new model architecture to llama.cpp.

    149 GitHub stars~4.1k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Turbofit

    SouthpawIN/turbofit

    Hardware-aware adaptive Hermes runtime using portable Turbofiles, total usable memory, owned native llama.cpp residency, stable auto/active:main/active:aux routes, and evidence-backed promotion.

    107 GitHub stars~1.9k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed

More from scragnog/HOT-Step-CPP

All 18 skills in this repo
  • Ear Test Scoresheet

    scragnog/HOT-Step-CPP

    The standard way to run a listening test in HOT-Step - a local HTML score sheet next to the renders where Rob plays each track, scores it 1-5 on named criteria, and the page charts the two score…

    173 GitHub stars~1.9k tokensUpdated 2 days ago
    Auto-check passed
  • Engine Performance

    scragnog/HOT-Step-CPP

    Explains where HOT-Step generation time goes (LM/DiT/VAE), how the TensorRT paths activate, how to benchmark from logs, and which knobs trade quality for speed.

    173 GitHub stars~4.9k tokensUpdated 2 days ago
    Auto-check passed
  • Mm3 Backend

    scragnog/HOT-Step-CPP

    Maps HOT-Step's native MiniMax-Music3 backend — engine port modules, endpoints, server/UI integration, parity/fixture infrastructure, and the hard-won trap list.

    173 GitHub stars~4.5k tokensUpdated 2 days ago
    Auto-check passed
  • Mm3 Lm Adapter Training

    scragnog/HOT-Step-CPP

    The validated recipe for training MiniMax-Music3 planner-LM style adapters (artist/album clones) with ace-train mm3-lm-train and the Training Studio.

    173 GitHub stars~4k tokensUpdated 2 days ago
    Auto-check passed
  • Release Process

    scragnog/HOT-Step-CPP

    Runbook for cutting and publishing a HOT-Step CPP release via a v git tag that triggers the multi-platform CI build and drafts a GitHub Release.

    173 GitHub stars~5.1k tokensUpdated 2 days ago
    Auto-check passed
  • Upstream Sync

    scragnog/HOT-Step-CPP

    Safely pulls upstream acestep.cpp changes into the HOT-Step engine fork without destroying its integration hooks.

    173 GitHub stars~5k tokensUpdated 2 days ago
    Auto-check passed

Questions about Model Management

What does Model Management do?

Explains how model files, checkpoints, GGUF quantization, and the Model Manager work in HOT-Step CPP. Model Management is an agent skill from scragnog/HOT-Step-CPP. Explains how model files, checkpoints, GGUF quantization, and the Model Manager work in HOT-Step CPP.

When should I use Model Management?

Model Management fits situations like: adding/converting/quantizing models; debugging no GGUF models found; missing-model/wrong-model failures; working on the model download service.

How do I install Model Management in Claude Code?

Run `npx skills add scragnog/HOT-Step-CPP --skill model-management -a claude-code`. Or copy the skill folder (.claude/skills/model-management in scragnog/HOT-Step-CPP) into .claude/skills/model-management in your project. Claude Code loads it when a task matches its description.

How do I install Model Management in Codex?

Run `npx skills add scragnog/HOT-Step-CPP --skill model-management -a codex`. Or copy the skill folder (.claude/skills/model-management in scragnog/HOT-Step-CPP) into .agents/skills/model-management in your project. Codex loads it when a task matches its description.

Can I use Model Management in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add scragnog/HOT-Step-CPP --skill model-management -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/model-management, .gemini/skills/model-management, .github/skills/model-management and .opencode/skills/model-management in your project.

What does Model Management need to run?

Going by SKILL.md and its folder, Model Management needs the command-line tools its instructions call (node, python, cmake and pip). Our summary lists: Python 3.

Does Model Management access the network?

SKILL.md names 1 domain. In commands or code: huggingface.co; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Model Management safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Model Management use?

Model Management is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Model Management use?

About 6.6k tokens (SKILL.md is roughly 26k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Model Management?

Skills that share tags, products or a category with Model Management: Qwen Mtp Gguf (R6410418/Jackrong-llm-finetuning-guide, 1.7k stars), Add Model (guoqingbao/xinfer, 334 stars), Resolve (alexziskind1/model-shelf, 130 stars) and Test Model (guoqingbao/xinfer, 334 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Model Management?

scragnog (a GitHub user) maintains it in scragnog/HOT-Step-CPP, which has 173 GitHub stars. The repository holds 18 skills in this directory. The repository was last updated on October 7, 2026.

Source: scragnog/HOT-Step-CPP on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.