Extempore JIT Debugging Guide
digego/extempore
Debugging guide for Extempore covering its three layers, compilation paths, startup sequence and the batch, eval and interactive modes used to isolate JIT problems.
Diagnoses HOT-Step CPP generation failures, engine crashes, hangs, and startup problems from the logs/ session folders.
$ npx skills add scragnog/HOT-Step-CPP --skill debugging-runtime -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install scragnog/HOT-Step-CPP debugging-runtime --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/scragnog/HOT-Step-CPP.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/debugging-runtime .claude/skills/debugging-runtime && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "debugging-runtime" agent skill from https://github.com/scragnog/HOT-Step-CPP/tree/master/.claude/skills/debugging-runtime into .claude/skills/debugging-runtime/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "debugging-runtime", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/scragnog/HOT-Step-CPP/tree/master/.claude/skills/debugging-runtimeType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add scragnog/HOT-Step-CPP --skill debugging-runtime -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install scragnog/HOT-Step-CPP debugging-runtime --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/scragnog/HOT-Step-CPP.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills/debugging-runtime .agents/skills/debugging-runtime && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "debugging-runtime" agent skill from https://github.com/scragnog/HOT-Step-CPP/tree/master/.claude/skills/debugging-runtime into .agents/skills/debugging-runtime/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "debugging-runtime", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add scragnog/HOT-Step-CPP --skill debugging-runtime -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install scragnog/HOT-Step-CPP debugging-runtime --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/scragnog/HOT-Step-CPP.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills/debugging-runtime .cursor/skills/debugging-runtime && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "debugging-runtime" agent skill from https://github.com/scragnog/HOT-Step-CPP/tree/master/.claude/skills/debugging-runtime into .cursor/skills/debugging-runtime/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "debugging-runtime", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/scragnog/HOT-Step-CPP.git --path .claude/skills/debugging-runtime--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add scragnog/HOT-Step-CPP --skill debugging-runtime -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install scragnog/HOT-Step-CPP debugging-runtime --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/scragnog/HOT-Step-CPP.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills/debugging-runtime .gemini/skills/debugging-runtime && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "debugging-runtime" agent skill from https://github.com/scragnog/HOT-Step-CPP/tree/master/.claude/skills/debugging-runtime into .gemini/skills/debugging-runtime/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "debugging-runtime", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install scragnog/HOT-Step-CPP debugging-runtimeInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add scragnog/HOT-Step-CPP --skill debugging-runtime -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/scragnog/HOT-Step-CPP.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills/debugging-runtime .github/skills/debugging-runtime && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "debugging-runtime" agent skill from https://github.com/scragnog/HOT-Step-CPP/tree/master/.claude/skills/debugging-runtime into .github/skills/debugging-runtime/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "debugging-runtime", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add scragnog/HOT-Step-CPP --skill debugging-runtime -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install scragnog/HOT-Step-CPP debugging-runtime --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/scragnog/HOT-Step-CPP.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills/debugging-runtime .opencode/skills/debugging-runtime && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "debugging-runtime" agent skill from https://github.com/scragnog/HOT-Step-CPP/tree/master/.claude/skills/debugging-runtime into .opencode/skills/debugging-runtime/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "debugging-runtime", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
debugging-runtimeDiagnoses HOT-Step CPP generation failures, engine crashes, hangs, and startup problems from the logs/ session folders.
Debugging Runtime is an agent skill from scragnog/HOT-Step-CPP. Diagnoses HOT-Step CPP generation failures, engine crashes, hangs, and startup problems from the logs/ session folders. Use when a music generation failed, ace-server crashed or keeps respawning, the app hangs with no progress, an API call returns 500/503, or you need to trace a gen<uuid log back to engine output.
Its SKILL.md is about 6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `reference.md`).
It sits in Media & Creative, covering Music and audio generation and Debugging. It works with C++. The repository describes itself as: Turn dials. Summon bangers! NOW WITH MORE C++! Local AI music generation powered by GGML. The licence is MIT.
7 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 24b12b5. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
cmakenodenpxFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use npx, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Debugging Runtime loads about 6k tokens when it runs. Until then it costs about 84 tokens; SKILL.md has 2,780 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
-52); override with `ACESTEPCPP_EXE` in `.env`.Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from scragnog/HOT-Step-CPP at commit 24b12b5, republished under its MIT licence (© scragnog). 2,780 words, ~5,969 tokens.
.claude/skills/debugging-runtime/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.HOT-Step CPP has two runtime processes: the Node server (Express, port 3001) and its child ace-server.exe (the C++ inference engine, port 8085). Agents work in dev mode (dev.bat), where the Vite dev server on 3000 fronts the app and proxies /api, /audio and /references to 3001 — every URL below is the 3000 address. The Node server orchestrates every generation: it optionally calls the engine's LM (language model that expands your caption/lyrics into audio codes), then submits synthesis (DiT — the Diffusion Transformer that generates audio latents — followed by VAE decode to a WAV), polls until done, then saves the result to SQLite. Almost every runtime problem is diagnosable from the per-session log folders under logs/ at the repo root. This skill tells you which file to open first for each symptom, what the real failure strings mean, and how to correlate the three log files.
All path:line references verified against the code on 2026-07-02.
POST /api/generate returns 503 "Engine not ready", or any API route 500s.taskkill / Stop-Process) while the Node server is running. The Node server auto-respawns the engine on non-zero exit (server/src/index.ts:284-309); if crashes are spaced more than 30 s apart, the crash limiter's window resets and the respawn loop continues indefinitely, holding file locks on the exe. To rebuild the engine, use dev-rebuild.bat at the repo root — it shuts the whole app down cleanly first, then builds. To just stop the engine, use Invoke-RestMethod -Method Post http://localhost:3000/api/shutdown (kills engine + server + Vite).engine/build.cmd directly, under any circumstances — you cannot reliably tell whether the app is running; same respawn/file-lock reason. dev-rebuild.bat wraps it safely (and is a harmless no-op shutdown when nothing runs).cmake --build . --clean-first — CUDA kernel recompilation takes 20+ minutes. For stale .obj problems, delete only engine/build/acestep-core.dir/ and engine/build/Release/acestep-core.lib.ace_engine.log for advancing [DiT] Step N/M lines before declaring it hung.;, never &&.One folder per Node-server session, named so name-sort = time-sort. Created at server boot by initLogger() (server/src/services/logger.ts:40):
logs/YYYY-MM-DD_HH-MM-SS/
node_console.log all Node stdout+stderr, mirrored transparently
ace_engine.log raw ace-server child stdout+stderr
generations/
gen_<jobId>_<taskType>.log one per generation; jobId = server-side UUID
training/
<kind>_<jobId>.log one per training job, written live: `> $ <trainer command>`,
every job log line (INFO/WARN/ERROR), the trainer's raw
stdout/stderr (`> ...`), and `END | <status>`Facts you must know before reading them:
node_console.log contains every engine line EXCEPT the filtered GGML noise patterns (which appear only in ace_engine.log — index.ts:262-272 gates console output on isNoise() but writes ace_engine.log unconditionally). Engine lines appear prefixed [ace-server] , interleaved with server lines in real arrival order — it is the only file where engine and server output are time-ordered relative to each other.ace_engine.log and node_console.log have NO per-line timestamps. Only gen_*.log lines carry ISO timestamps (2026-07-01T10:28:33.514Z | INFO | ..., logger.ts:113-119).gen_*.log is buffered in RAM and only written to disk when the generation completes or fails (finishGenerationLog / failGenerationLog, logger.ts:139-178 — both do a single fs.writeFileSync). Failed and cancelled generations DO get their log written. If Node itself crashes or is hard-killed mid-generation, the gen log is never written. A missing gen_*.log for a generation you know started = Node died mid-flight; fall back to node_console.log.<taskType> in the filename comes from the engine request's task_type: text2music (default), cover, cover-nofsq, repaint, lego, extract (generate.ts:208-213). The retry-exhausted final-failure path writes taskType 'unknown' (generate.ts:1335), but the earlier per-attempt failure (generate.ts:1275/1281) already flushed and deleted the buffer, so gen_<id>_unknown.log is usually a silent no-op.CUDA graph warmup, CUDA Graph id, ggml_backend_cuda_graph_compute) is dropped at the engine source since 2026-07-17: acestep_ggml_log (engine/src/backend.h) discards all GGML DEBUG-level messages (set HOTSTEP_GGML_DEBUG=1 to pass them through) and digit-insensitively dedups consecutive near-identical lines, so these no longer reach ANY log file. The Node-side filters (index.ts isNoise(), server/src/routes/logs.ts:27-31) remain as belt-and-braces for older engine binaries. If you need CUDA-graph-layer logging, use the env var.GET http://localhost:3000/api/logs is an SSE stream backed by a 2000-line ring buffer, each line tagged source: 'engine' | 'server' (logs.ts:21-52).| Symptom | Open first | Then | Looking for |
|---|---|---|---|
| Generation failed (UI error) | Newest logs/<session>/generations/gen_<uuid>_*.log — last line is GENERATION FAILED: <reason> | node_console.log around that job; ace_engine.log for C++ detail | The failure reason string (table below) |
| Engine crash | node_console.log | ace_engine.log tail | [ace-server] Process exited with code N + the FATAL/assert lines just before it |
| Server 500 / API error | node_console.log | — | Express stack traces; [Server] Uncaught exception / Unhandled rejection (logged and swallowed — process keeps running, index.ts:576-581) |
| Startup failure | node --version FIRST, then node_console.log | ace_engine.log | Node 20 to 24 LTS, 24 recommended (engines in server/package.json enforces <25; use matching native modules). Then: [Server] ace-server not found at:, CUDA-runtime download banner, Crashed 3 times within 30s — giving up, DB errors |
| Hang / no progress | GET /api/generate/queue (live), then node_console.log tail | ace_engine.log tail | Last engine line = the wedged phase. The 120 s stall watchdog usually converts hangs into a Generation stalled failure on its own |
| No gen log exists at all | node_console.log | — | Node died mid-generation (gen logs only flush at the end) |
No session folder at all → the server never reached initLogger(); run npx tsx src/index.ts from server/ and read the terminal directly.
$s = (Get-ChildItem logs | Sort-Object Name -Descending | Select-Object -First 1).FullName
Get-Content "$s\node_console.log" -Tail 80
Get-Content "$s\ace_engine.log" -Tail 60
Get-ChildItem "$s\generations" | Sort-Object LastWriteTime -Descending | Select-Object -First 3GENERATION FAILED: <reason>; the top of the file embeds the full resolved engine request JSON (seed, adapters, solver, scheduler — logger.ts:124-133). That JSON is ground truth for reproduction.Generation failed on ace-server), the real cause is engine-side only — grep the session:Select-String -Path "$s\node_console.log" -Pattern "exited with code|Crashed|FATAL|CUDA error"Invoke-RestMethod http://localhost:3000/api/health | ConvertTo-Json -Depth 4
Invoke-RestMethod http://localhost:3000/api/generate/queue/api/health reports aceServer.status (ok/disconnected) and engine.{ready, bootStatus} (server/src/routes/health.ts:11-46). Remember rule 4: a disconnected aceServer during heavy compute can be a busy single-threaded engine, not a dead one.Invoke-RestMethod -Method Post http://localhost:3000/api/generate/reset-queueQueue reset by user) and drains the pending queue (generate.ts:1469-1500). Also available: POST /api/generate/cancel/:id, POST /api/generate/cancel-all.Select-String -Path "logs\*\generations\*.log" -Pattern "GENERATION FAILED" | Select-Object -Last 10GENERATION FAILED: <X>| Signature | Cause / next step |
|---|---|
Cancelled by user | Benign — user hit cancel (generate.ts:1275). Most common "failure" in history. |
Generation stalled — no progress for Ns (last stage: "...") | The 120 s stall watchdog fired (generate.ts:128-134). The quoted stage names the wedged phase (e.g. "Loading CondEnc..."); check ace_engine.log tail for what the engine was doing. Also fires when the engine died silently and polls just time out. |
Generation timed out (N min limit) | Wall-clock watchdog: default 45 min, user-clamped 5–120 via generationTimeoutMinutes (generate.ts:105-106, 139). Huge duration/steps, or a first-run TensorRT engine build without a cached engine. |
Generation failed on ace-server | Engine set job status failed without crashing (generate.ts:156). Deliberately generic — the real reason is only in ace_engine.log; grep it for FATAL, [Server], failed. |
...POST /synth... failed (413) | Engine rejected the upload: audio exceeds max duration (10 min) or src_latents exceeds max frames (engine/tools/hot-step-server.cpp:2044, 2072); global payload cap is 256 MB (hot-step-server.cpp:2700). Source audio for a cover/repaint is too long. |
Engine not ready: <status> (HTTP 503, never reaches a gen log) | Bootstrap incomplete (CUDA DLL download still running) or the crash limiter gave up (generate.ts:1356-1363). Check node_console.log. |
Queue reset by user | Someone hit /api/generate/reset-queue. Benign. |
ace_engine.log / [ace-server] lines)| Signature | Cause |
|---|---|
[VAE] FATAL: tensor 'decoder.conv1.weight_v' not found in safetensors then process exit | Wrong or corrupt VAE model file (real occurrence 2026-06-04). |
[Server] FATAL: LM load failed / [Server] FATAL: synth load failed | Model file missing/corrupt or OOM during load (hot-step-server.cpp:831, 1232). Job goes to failed; the engine process usually survives. |
[Safetensors] Cannot open <path>/adapter_model.safetensors | Usually benign filename probing — the adapter loader tries candidate filenames. Thousands appear in healthy sessions. Only meaningful if the adapter then actually fails to load. |
CUDA error: ... (any) | GGML/CUDA fault; typically followed by process exit and respawn. |
Process exited with code 3221225786 (0xC000013A) | Console Ctrl+C / window closed — usually deliberate. |
Process exited with code 3221226505 (0xC0000409) | Fail-fast / abort / stack-buffer-overrun — a genuine C++ crash. Get the engine lines immediately preceding it. |
Process exited with code 1, signal null right after [Shutdown] Killed ace-server PID ... or [Server] Shutting down... | Benign — taskkill /F yields exit code 1, not a signal. Only treat an exit line as a crash if it is NOT preceded by a shutdown line. |
[ace-server] Restarting in 3 seconds... (crash N/3) | Routine supervised respawn — dozens exist in healthy log history. Investigate the crash cause, not the respawn itself. |
[ace-server] Crashed 3 times within 30s — giving up. | Crash limiter tripped (index.ts:296-301). Usually missing DLLs (cuBLAS runtime on portable CUDA builds) — reconnect to the internet and restart, or re-extract the release zip. engineReady becomes false; /api/generate returns 503. |
| All solvers/schedulers/guidance behave identically or ignore settings, logs completely clean — especially after an upstream acestep.cpp sync | The hot-step-sampler.h hook in pipeline-synth-ops.cpp was silently overwritten (compiles fine, all plugins dead — invisible in logs). Run engine/verify-hooks.ps1; see the upstream-sync skill. |
startAceServer() (index.ts:158-316) launches config.aceServer.exe with --models, --host, --port (default 8085, server/src/config.ts:102) plus optional --adapters, --keep-loaded, --noise-profile, --draft-lm, --vae-chunk, --vae-overlap. The exe is auto-detected among engine/ace-server.exe, engine/build/Release/ace-server.exe, engine/build/ace-server.exe, engine/build/Debug/ace-server.exe (config.ts:46-52); override with ACESTEPCPP_EXE in .env.engineReady=false with bootStatus Engine crashed 3 times — check logs for missing DLLs (index.ts:152-156, 284-309). Crashes spaced >30 s apart reset the window, so slow-cycle crashes (e.g. crash-on-first-request) still loop forever — hence Golden rule 1.shutdown() uses taskkill /PID <child> /T /F (index.ts:549). POST /api/shutdown kills ace-server by port 8085 via netstat (server/src/routes/shutdown.ts:24-44), then Vite (:3000), then its own process tree. POST /api/restart writes a .restart-requested marker (shutdown.ts:161); in dev mode server/restart-loop.cmd loops on it and relaunches Node (LAUNCH.bat is the end-user prod equivalent).dev-rebuild.bat = POST /api/shutdown → wait up to 10 s for ace-server.exe to die (force-kill at 10 s, abort at 15 s) → engine\build.cmd. It does not restart the app — run dev.bat afterwards (agents never use LAUNCH.bat).POST /api/generate returns 503 Engine not ready: <bootStatus>. No internet → CPU-only start with a banner in node_console.log.enqueueGeneration, generate.ts:1291-1352), because engine log parsing is a global untagged pub/sub.enqueueGeneration has a retry-once-with-fresh-seed loop (generate.ts:1301-1339), but runGeneration's top-level catch (generate.ts:1271-1283) swallows every generation error without rethrowing, so the retry never fires (the only escapable throw, translateParams at :191, happens before any engine submission). One UI-visible failure = exactly ONE engine attempt, and the seed in the gen log's single params block is directly trustworthy for repro. Zero [Retry] lines exist across 118 logged sessions. If retry is ever intended behavior, the swallowed rethrow is the server bug to fix.pollUntilDone (generate.ts:102-170) polls the engine every 500 ms; stage/progress unchanged for 120 s → cancel + Generation stalled. First-run TensorRT builds are special-cased ([TRT-WARN] lines mutate the stage string, generate.ts:756) so a 5–10 min TRT compile doesn't trip it.[Generate] Poll error ... (will retry) (generate.ts:165) just means the busy single-threaded engine missed a 30 s poll window.GET /api/generate/status/:id as ace_phase — one of queued, loading_text_enc, encoding_text, loading_cond_enc, encoding_cond, loading_dit, loading_adapter, adapter_precompute, dit_inference, loading_vae, vae_decode, encoding_output, done, failed, cancelled (aceClient.ts:165-180) — plus ace_phase_progress ("step N/M"). These fields are optional on the wire for older engine builds; absence is not an error./status/:id on an old job is TTL cleanup, not data loss — the song row and gen log persist.WARNING and the job still succeeds with raw audio. A "succeeded" job can still have WARNINGs worth reading in its gen log.<uuid> log with engine + node logsThere are two job-ID namespaces: the server UUID (the gen_<uuid> filename; SSE lines show only its first 8 chars as [Gen:xxxxxxxx], logger.ts:118) and the engine's own job id (ace_job_id in /status/:id; returned by the engine's /lm and /synth). The engine id is NOT in the gen filename.
gen_<uuid>_*.log; note the ISO timestamps of the LM/synth submission lines and the final GENERATION FAILED line.ace_engine.log has no timestamps, so pivot through node_console.log: search for [Generate] Job <uuid> — the submission line logs ditModel/synth_model/seed/source (generate.ts:199) and the failure logs [Generate] Job <uuid> failed: (generate.ts:1280). The block of [ace-server] ... lines between your job's submit and fail lines IS that generation's engine output.ace_engine.log (e.g. to see noise-filtered lines), grep it for a distinctive engine line found in step 2 (the last [DiT] Step N/M, a model-load line) and read forward from there.Dev-mode note: under dev.bat (tsx watch), every source-change auto-restart creates a new session folder. A burst of near-identical logs/ folders seconds apart means tsx was restarting, not that the app was crashing.
| Path | Role |
|---|---|
server/src/services/logger.ts | Session folders, console/engine mirroring, gen-log buffering & flush |
server/src/index.ts | Engine spawn, respawn + crash limiter (:152-316), shutdown (:540), CUDA DLL bootstrap |
server/src/routes/generate.ts | Orchestration: pollUntilDone :102, runGeneration :173, retry/queue :1291, status/queue/reset endpoints :1390+ |
server/src/services/aceClient.ts | Typed engine HTTP client; timeouts; AceJobPhase list |
server/src/routes/logs.ts | GET /api/logs SSE, 2000-line ring buffer, noise filter |
server/src/routes/health.ts | GET /api/health |
server/src/routes/shutdown.ts | POST /api/shutdown (port-based kill), POST /api/restart marker |
server/src/engineState.ts | engineReady / engineBootStatus gate |
server/src/config.ts | Engine exe auto-detect, port 8085, env-var overrides |
engine/tools/hot-step-server.cpp | C++ engine HTTP server; FATAL emit sites; 413/payload limits |
dev-rebuild.bat | The ONLY sanctioned way to rebuild the engine, always |
dev.bat | The launcher agents use: dev mode (Vite :3000 HMR + tsx watch on :3001), started detached |
LAUNCH.bat | End-user prod launcher (Node :3001 serving ui/dist/) — agents don't run it |
dev-rebuild.bat.ace_engine.log and node_console.log carry no per-line timestamps; only gen logs do. All time correlation goes through node_console.log line ordering.Process exited with code 1, signal null after a shutdown line is a taskkill artifact, not a crash (dozens in history). Respawns themselves are routine (Restarting in 3 seconds... (crash N/3) appears across many healthy sessions).[Safetensors] Cannot open ...adapter_model.safetensors is usually benign filename probing.enqueueGeneration is unreachable for generation failures (runGeneration swallows its own errors, generate.ts:1271-1283); no [Retry] line has ever appeared in log history. The gen-log params block is the real, only attempt.status:'failed' for every failure class — hot-step-server.cpp has many FATAL/failed emit sites; the tables above list only those with real log occurrences. Linux/macOS shutdown paths exist in code (shutdown.ts:45+) but are unconfirmed by logs. gen_<id>_unknown.log files have a code path but none have been observed on disk.--draft-lm is only passed if ACESTEPCPP_DRAFT_LM is set. Don't chase it as a crash suspect unless that env var is set.docs/plans/ — internal investigation docs. Gitignored, local-only — may be absent on a fresh clone.© scragnog, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 1 other file in .claude/skills/debugging-runtime of scragnog/HOT-Step-CPP.
Open the folder on GitHubat commit 24b12b5
Debugging Runtime next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Debugging Runtime this skillscragnog/HOT-Step-CPP | 170 | — | ~6k | Automated safety check: Notes | MIT | |
| Extempore JIT Debugging Guidedigego/extempore | 1.5k | — | ~4.4k | Automated safety check: Pass | None | |
| OpenROAD Bug FixerThe-OpenROAD-Project/OpenROAD | 3.2k | — | ~784 | Automated safety check: Pass | BSD-3-Clause | |
| Cppcrazyguitar/cppcheatsheet | 290 | — | ~1.8k | Automated safety check: Pass | MIT | |
| MCP Debuggerdebugmcp/mcp-debugger | 171 | — | ~3.8k | Automated safety check: Pass | MIT | |
| Dbgtheodo-group/debug-that | 158 | — | ~1.9k | Automated safety check: Pass | MIT |
digego/extempore
Debugging guide for Extempore covering its three layers, compilation paths, startup sequence and the batch, eval and interactive modes used to isolate JIT problems.
The-OpenROAD-Project/OpenROAD
Fixes an OpenROAD bug from a GitHub issue or error code: finds the root cause, implements the fix, adds a regression test and prepares a signed-off commit.
crazyguitar/cppcheatsheet
Comprehensive C/C++ programming reference covering everything from C11-C23 and C++11-C++23, system programming, CUDA GPU computing, debugging tools, Rust interop, and advanced topics.
debugmcp/mcp-debugger
A skill your agent uses when investigating a bug, failing test, or unexpected runtime behavior and the mcp-debugger MCP server is available — drives real step-through debuggers (breakpoints, stack…
theodo-group/debug-that
Debug applications using the dbg CLI debugger. An agent skill from theodo-group/debug-that.
vipshop/cache-dit
A skill your agent uses when writing, debugging, porting, reviewing, or optimizing CUDA C++ or PTX kernels; investigating CUDA Runtime or Driver API behavior; profiling kernels with Nsight Systems…
scragnog/HOT-Step-CPP
The standard way to run a listening test in HOT-Step - a local HTML score sheet next to the renders where Rob plays each track, scores it 1-5 on named criteria, and the page charts the two score…
scragnog/HOT-Step-CPP
Explains where HOT-Step generation time goes (LM/DiT/VAE), how the TensorRT paths activate, how to benchmark from logs, and which knobs trade quality for speed.
scragnog/HOT-Step-CPP
Maps HOT-Step's native MiniMax-Music3 backend — engine port modules, endpoints, server/UI integration, parity/fixture infrastructure, and the hard-won trap list.
scragnog/HOT-Step-CPP
The validated recipe for training MiniMax-Music3 planner-LM style adapters (artist/album clones) with ace-train mm3-lm-train and the Training Studio.
scragnog/HOT-Step-CPP
Runbook for cutting and publishing a HOT-Step CPP release via a v git tag that triggers the multi-platform CI build and drafts a GitHub Release.
scragnog/HOT-Step-CPP
Safely pulls upstream acestep.cpp changes into the HOT-Step engine fork without destroying its integration hooks.
Works with
Categories
Diagnoses HOT-Step CPP generation failures, engine crashes, hangs, and startup problems from the logs/ session folders. Debugging Runtime is an agent skill from scragnog/HOT-Step-CPP. Diagnoses HOT-Step CPP generation failures, engine crashes, hangs, and startup problems from the logs/ session folders.
Debugging Runtime fits situations like: A music generation failed; ace-server crashed; keeps respawning; the app hangs with no progress.
Run `npx skills add scragnog/HOT-Step-CPP --skill debugging-runtime -a claude-code`. Or copy the skill folder (.claude/skills/debugging-runtime in scragnog/HOT-Step-CPP) into .claude/skills/debugging-runtime in your project. Claude Code loads it when a task matches its description.
Run `npx skills add scragnog/HOT-Step-CPP --skill debugging-runtime -a codex`. Or copy the skill folder (.claude/skills/debugging-runtime in scragnog/HOT-Step-CPP) into .agents/skills/debugging-runtime in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add scragnog/HOT-Step-CPP --skill debugging-runtime -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/debugging-runtime, .gemini/skills/debugging-runtime, .github/skills/debugging-runtime and .opencode/skills/debugging-runtime in your project.
Going by SKILL.md and its folder, Debugging Runtime needs the command-line tools its instructions call (cmake, node and npx). Our summary lists: Node.js.
SKILL.md contains no URLs. Its commands use npx, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.
Debugging Runtime is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 6k tokens (SKILL.md is roughly 24k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Debugging Runtime: Extempore JIT Debugging Guide (digego/extempore, 1.5k stars), OpenROAD Bug Fixer (The-OpenROAD-Project/OpenROAD, 3.2k stars), Cpp (crazyguitar/cppcheatsheet, 290 stars) and MCP Debugger (debugmcp/mcp-debugger, 171 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
scragnog (a GitHub user) maintains it in scragnog/HOT-Step-CPP, which has 170 GitHub stars. The repository holds 18 skills in this directory. The repository was last updated on October 5, 2026.
Source: scragnog/HOT-Step-CPP on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.