Aristotle
K-Dense-AI/mimeographs
Applies the frameworks of Aristotle (ancient Greek philosopher, logic, ethics, metaphysics, 384-322 BCE) to decision-making, ethics, and analysis.
Diagnose and fix 3DGS training-time failures: NaN losses, OOM crashes, divergent optimization, floater artifacts, densification failures, hyperparameter sensitivity.
$ npx skills add jaccen/Awesome-Gaussian-Skills --skill 3dgs-training-debugger -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install jaccen/Awesome-Gaussian-Skills 3dgs-training-debugger --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/jaccen/Awesome-Gaussian-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/3dgs-training-debugger .claude/skills/3dgs-training-debugger && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "3dgs-training-debugger" agent skill from https://github.com/jaccen/Awesome-Gaussian-Skills/tree/main/skills/3dgs-training-debugger into .claude/skills/3dgs-training-debugger/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "3dgs-training-debugger", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/jaccen/Awesome-Gaussian-Skills/tree/main/skills/3dgs-training-debuggerType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add jaccen/Awesome-Gaussian-Skills --skill 3dgs-training-debugger -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install jaccen/Awesome-Gaussian-Skills 3dgs-training-debugger --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaccen/Awesome-Gaussian-Skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/3dgs-training-debugger .agents/skills/3dgs-training-debugger && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "3dgs-training-debugger" agent skill from https://github.com/jaccen/Awesome-Gaussian-Skills/tree/main/skills/3dgs-training-debugger into .agents/skills/3dgs-training-debugger/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "3dgs-training-debugger", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jaccen/Awesome-Gaussian-Skills --skill 3dgs-training-debugger -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install jaccen/Awesome-Gaussian-Skills 3dgs-training-debugger --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaccen/Awesome-Gaussian-Skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/3dgs-training-debugger .cursor/skills/3dgs-training-debugger && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "3dgs-training-debugger" agent skill from https://github.com/jaccen/Awesome-Gaussian-Skills/tree/main/skills/3dgs-training-debugger into .cursor/skills/3dgs-training-debugger/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "3dgs-training-debugger", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/jaccen/Awesome-Gaussian-Skills.git --path skills/3dgs-training-debugger--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add jaccen/Awesome-Gaussian-Skills --skill 3dgs-training-debugger -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install jaccen/Awesome-Gaussian-Skills 3dgs-training-debugger --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaccen/Awesome-Gaussian-Skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/3dgs-training-debugger .gemini/skills/3dgs-training-debugger && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "3dgs-training-debugger" agent skill from https://github.com/jaccen/Awesome-Gaussian-Skills/tree/main/skills/3dgs-training-debugger into .gemini/skills/3dgs-training-debugger/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "3dgs-training-debugger", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install jaccen/Awesome-Gaussian-Skills 3dgs-training-debuggerInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add jaccen/Awesome-Gaussian-Skills --skill 3dgs-training-debugger -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/jaccen/Awesome-Gaussian-Skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/3dgs-training-debugger .github/skills/3dgs-training-debugger && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "3dgs-training-debugger" agent skill from https://github.com/jaccen/Awesome-Gaussian-Skills/tree/main/skills/3dgs-training-debugger into .github/skills/3dgs-training-debugger/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "3dgs-training-debugger", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jaccen/Awesome-Gaussian-Skills --skill 3dgs-training-debugger -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install jaccen/Awesome-Gaussian-Skills 3dgs-training-debugger --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaccen/Awesome-Gaussian-Skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/3dgs-training-debugger .opencode/skills/3dgs-training-debugger && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "3dgs-training-debugger" agent skill from https://github.com/jaccen/Awesome-Gaussian-Skills/tree/main/skills/3dgs-training-debugger into .opencode/skills/3dgs-training-debugger/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "3dgs-training-debugger", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
3dgs-training-debuggerDiagnose and fix 3DGS training-time failures: NaN losses, OOM crashes, divergent optimization, floater artifacts, densification failures, hyperparameter sensitivity.
3dgs Training Debugger is an agent skill from jaccen/Awesome-Gaussian-Skills. Diagnose and fix 3DGS training-time failures: NaN losses, OOM crashes, divergent optimization, floater artifacts, densification failures, hyperparameter sensitivity. Covers runtime debugging for vanilla 3DGS and 50+ novel methods (deformable, MoE, physics-based, feed-forward). Detects 60 runtime failure patterns. Use when: 3DGS training crashes or produces poor results, loss is NaN/Inf, VRAM exhaustion, Gaussians explode or vanish, densification not working, convergence stalls, 训练调试/显存溢出/训练发散/浮点伪影.
Its SKILL.md is about 7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including reference files (for example `references/convergence-trajectories.md`, `references/runtime-bug-patterns.md` and `references/vram-gpu-table.md`).
It sits in Development, covering Creative writing and fiction. The repository describes itself as: 图形学与3DGS、空间智能持续更新论文;AI Agent Skills for 3D Gaussian Splatting, NeRF & Computer Graphics Research. 800+ methods, 25categories, 12skills. OpenClaw / Claude Code compatible. The licence is Apache-2.0.
6 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 437c820. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are python).
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
3dgs Training Debugger loads about 7k tokens when it runs, and up to ~19k if it reads all its reference files. Until then it costs about 132 tokens; SKILL.md has 2,522 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from jaccen/Awesome-Gaussian-Skills at commit 437c820, republished under its Apache-2.0 licence (© jaccen). 2,522 words, ~7,003 tokens.
.claude/skills/3dgs-training-debugger/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.You are a senior 3DGS engineer who has trained hundreds of Gaussian Splatting models across vanilla 3DGS, deformable GS, feed-forward GS, SLAM-GS, and physics-based GS pipelines. Diagnose and fix training-time failures systematically.
This skill covers the runtime training phase — what happens AFTER code is written and BEFORE evaluation. It complements:
| Metric | Expected Behavior | Alert Threshold | Log Frequency |
|---|---|---|---|
| L1 loss | Decreasing, minor oscillation | Increase > 20% over 500 iters | Every 50 iters |
| SSIM loss | Decreasing smoothly | Stagnant for 1000+ iters | Every 100 iters |
| Total loss | Decreasing, plateau ~70-80% of training | NaN, Inf, or sudden spike | Every 50 iters |
| PSNR (eval) | Increasing, plateau near end | Drop > 2dB between evals | Every 1000 iters |
| Gaussian count | Growth phase (0-15k), then stable | Explosive growth (>10x) or vanishing | Every 500 iters |
| VRAM usage | Stable with minor fluctuation during ADC | > 90% of total VRAM | Every 100 iters |
| Gradient norms | Stable, < 1.0 typically | > 10.0 or exactly 0.0 | Every 100 iters |
| Learning rate | Following schedule (warmup → cosine decay) | Unexpected reset or spike | Every 500 iters |
| Active Gaussians | Growth then pruning equilibrium | All pruned (count → 0) | Every ADC cycle |
| ADC trigger count | Periodic (every ~100 iters) | Never triggers or triggers every iter | Every ADC cycle |
Reference trajectory for vanilla 3DGS on a typical Mip-NeRF 360 scene:
Iter Loss PSNR Gaussians VRAM Notes
0 0.45 16.2 1(x SfM) 4.2 GB Init from SfM points
500 0.22 20.1 12,000 5.1 GB First ADC cycle
1000 0.15 23.5 45,000 7.8 GB Rapid growth phase
2000 0.09 26.8 180,000 12.4 GB Growth slowing
5000 0.05 29.2 350,000 14.2 GB Near convergence
7000 0.04 30.1 380,000 14.5 GB Fine-tuning
10000 0.03 30.4 390,000 14.6 GB Final (shutting densif)
15000 0.03 30.5 395,000 14.6 GB Densif frozen, opacity fine-tune
30000 0.025 30.6 395,000 14.6 GB Final modelKey signals: Gaussian count should plateau around iter 10k-15k (when densification freezes), PSNR should still improve slightly afterward via opacity/SH refinement.
See references/convergence-trajectories.md for expected trajectories across datasets, scene types, and method variants.
Start from the observed symptom and follow the branches.
SYMPTOM: Training crashed or NaN
│
├── NaN/Inf in loss?
│ ├── Check gradient norms → extremely large?
│ │ └── Possible: learning rate too high, gradient explosion
│ │ → Reduce lr to 1/10, add gradient clipping (max_norm=1.0)
│ ├── Check after ADC cycle → NaN appears right after densification?
│ │ └── Possible: NaN from clone/split, new Gaussian has bad scale/opacity
│ │ → Check scale clamping, opacity init values
│ ├── NaN from iteration 0?
│ │ └── Possible: bad initialization (zero covariance, SfM failure)
│ │ → Check point cloud, add covariance regularization
│ └── NaN in novel method (deformable/MoE)?
│ └── See Section 9: Novel Method Stability
│
├── CUDA OOM?
│ ├── During training (non-ADC)?
│ │ ├── Gaussian count reasonable but still OOM?
│ │ │ └── Possible: image resolution too high, batch size, SH degree
│ │ │ → Reduce image res by 2x, reduce SH to degree 1
│ │ └── Gaussian count exploding?
│ │ └── Possible: densification over-triggering
│ │ → See Section 3: OOM & Memory Management
│ └── During ADC (densification)?
│ └── Possible: temporary spike from clone/split
│ → Reduce ADC batch size, or move ADC to CPU
│
├── Training runs but quality is poor (low PSNR)?
│ ├── Gaussian count too low?
│ │ └── Possible: densification thresholds too strict, pruning too aggressive
│ │ → Lower grad_threshold, raise prune_threshold
│ ├── Gaussian count normal but artifacts?
│ │ └── See Section 4: Artifact Diagnosis
│ ├── Convergence stalled early?
│ │ └── See Section 6: Convergence Analysis
│ └── Specific views are bad?
│ └── Possible: training/test view selection issue, SfM sparse in that area
│
├── Training runs but visual artifacts?
│ ├── Floaters (small isolated Gaussians)?
│ │ └── See Pattern FP-01 in references
│ ├── Blur / over-smoothing?
│ │ └── See Pattern FP-02
│ ├── Ghosting / duplicate geometry?
│ │ └── See Pattern FP-03
│ ├── Holes / missing regions?
│ │ └── See Pattern FP-04
│ └── Color bleeding / SH artifacts?
│ └── See Pattern FP-05
│
└── Checkpoint resume gives different results?
└── See Section 8: Checkpoint & ResumeApproximate peak VRAM during training:
VRAM_peak ≈ Model_VRAM + Optimizer_VRAM + Raster_VRAM + Gradient_VRAM + ADC_spike
Where:
Model_VRAM = N_gaussians × bytes_per_gaussian
Optimizer_VRAM = 2 × Model_VRAM (Adam: momentum + variance)
Raster_VRAM = H × W × n_channels × num_images_in_batch × 4 bytes
Gradient_VRAM = Model_VRAM (gradients for all params)
ADC_spike = 1.5 × Model_VRAM (temporary allocation during clone/split)
bytes_per_gaussian ≈ 59 × 4 = 236 bytes
(3 position + 3 scale + 4 rotation + 1 opacity + 48 SH (degree 3) = 59 floats)See references/vram-gpu-table.md for precomputed VRAM requirements across Gaussian counts, SH degrees, and GPU types.
| Priority | Strategy | VRAM Savings | Quality Impact | Implementation |
|---|---|---|---|---|
| 1 | Reduce image resolution (2x downsample) | 50-75% raster VRAM | Minor PSNR drop (~0.5-1dB) | --data_factor 2 |
| 2 | Lower SH degree (3→1) | ~40% model VRAM | Slight view-dependent color loss | --sh_degree 1 |
| 3 | Gradient checkpointing on rasterizer | 30-40% gradient VRAM | ~10% slower training | Custom backward pass |
| 4 | Reduce ADC frequency (100→200 iters) | Reduces ADC spike frequency | Slower densification | --densify_interval 200 |
| 5 | CPU-offload optimizer states | 40% total VRAM | ~30% slower (PCIe transfer) | FSDP/DeepSpeed |
| 6 | Mixed precision (FP16/BF16) training | 30-50% total VRAM | Risk of numerical instability | torch.cuda.amp |
| 7 | Streaming image loading (not all in VRAM) | Major for large datasets | No quality impact | Custom data loader |
| 8 | Prune far-away Gaussians aggressively | Reduces model VRAM | May lose background detail | Custom prune criterion |
| Scenario | Typical Cause | Fix |
|---|---|---|
| OOM at iter ~500 (first ADC) | Sudden Gaussian count jump | Pre-allocate buffer for 5x initial count |
| OOM only on specific scenes | High-detail scenes grow more Gaussians | Scene-adaptive resolution reduction |
| OOM after checkpoint resume | Optimizer state not saved/loaded | Save full optimizer state in checkpoint |
| OOM on multi-GPU | All-reduce buffer too large | Gradient bucketing, overlap comm/compute |
| OOM with novel method | Extra params (deformation, MLP) | Profile each component separately |
| Artifact | Visual Symptom | Most Likely Training Cause | Diagnostic Action |
|---|---|---|---|
| Floaters | Small bright/dark blobs floating in space | Insufficient opacity pruning; ADC cloning noise | Check prune_opacity threshold; check if ADC ran after iter 15k |
| Blur | Overall soft, lacks high-freq detail | SH degree too low; low-resolution training images | Increase SH to 3; check --data_factor |
| Over-smoothing | PSNR OK but LPIPS bad, looks "flat" | L1+SSIM loss too weighted to L1; insufficient iterations | Increase SSIM weight (λ_dssim > 0.2) |
| Ghosting | Duplicate/semi-transparent geometry | Clone in wrong direction; scale gradient sign error | Check ADC clone position offset; verify gradient direction |
| Holes | Black/empty regions in reconstruction | Over-aggressive pruning; SfM sparse in that region | Raise prune threshold; add points in sparse areas |
| Color bleeding | Color from one surface leaks to another | SH coefficient overflow; insufficient view coverage | Clamp SH values; check training camera distribution |
| Stretching | Elongated Gaussian streaks | Scale not clamped; bad covariance projection | Verify scale_activation clamping (max 0.1-10.0) |
| Popping | View-dependent flickering between views | SH degree too high with sparse views; opacity reset | Reduce SH degree; increase opacity reset iterations |
| Z-fighting | Flickering on overlapping surfaces | Near-duplicate Gaussians at same depth | Add uniqueness in clone; increase prune threshold |
| Dark scene | Overall too dark / underexposed | Background color set to black; insufficient training | Set background to white or random; train longer |
Knowing WHEN the artifact was introduced narrows the cause:
Artifact present from iter 0 → Initialization issue (SfM points, scale init)
Artifact appears after first ADC → Densification bug (clone/split logic)
Artifact appears after 50% train → Pruning removed important Gaussians
Artifact appears near end → Opacity reset or SH overfitting
Artifact only in eval (not train) → Overfitting / view-dependent overfit| Parameter | Default | Range | Effect of Increase | Effect of Decrease |
|---|---|---|---|---|
position_lr | 0.00016 | 1e-5 to 1e-2 | Faster convergence, risk of explosion | Slower, more stable |
feature_lr | 0.0025 | 1e-4 to 1e-1 | Faster SH convergence | Slower color |
opacity_lr | 0.05 | 1e-3 to 0.2 | Faster opacity adaptation | Slower prune response |
scaling_lr | 0.005 | 1e-4 to 0.05 | Faster scale adaptation | More rigid geometry |
rotation_lr | 0.001 | 1e-5 to 0.01 | Faster rotation adaptation | More rigid orientation |
densify_grad_threshold | 0.0002 | 1e-5 to 1e-2 | More Gaussians (sensitive) | Fewer Gaussians |
densify_interval | 100 | 50-500 | Less frequent densification | More frequent |
densify_until_iter | 15000 | 5000-30000 | Longer growth phase | Earlier freeze |
prune_opacity_threshold | 0.005 | 0.001-0.05 | More aggressive pruning (fewer floaters) | More Gaussians (risk floaters) |
opacity_reset_interval | 3000 | 1000-10000 | More frequent resets (less view-dep overfit) | More stable opacities |
sh_degree | 3 | 0-4 | Better view-dependent color | Less VRAM |
lambda_dssim | 0.2 | 0-1 | More structural similarity | More pixel-level accuracy |
| Problem | First Adjustment | Second Adjustment | Last Resort |
|---|---|---|---|
| Low PSNR | Lower densify_grad_threshold (more Gaussians) | Increase training iterations | Lower image resolution |
| OOM | Lower sh_degree | Reduce image resolution | Decrease densify_until_iter |
| Floaters | Raise prune_opacity_threshold | Increase opacity_reset_interval | Post-train prune |
| Blur | Increase sh_degree | Increase lambda_dssim | Higher resolution images |
| Slow convergence | Increase position_lr | Increase densify_interval | Fewer total iters (accept lower quality) |
| Divergence | Decrease all LRs by 10x | Add gradient clipping | Reduce batch complexity |
| Too many Gaussians | Raise densify_grad_threshold | Lower densify_until_iter | Aggressive pruning |
Phase 1: Rapid Growth (iter 0 - 2,000)
- Loss drops fast, PSNR jumps from ~16 to ~24
- Gaussian count grows from SfM initial to ~50k-100k
- Risk: ADC over-triggering → OOM
Phase 2: Refinement (iter 2,000 - 15,000)
- Loss decreases more slowly, PSNR 24 → 28
- Gaussian growth slows, pruning starts balancing
- Risk: Premature densification freeze
Phase 3: Fine-tuning (iter 15,000 - 30,000)
- Loss near plateau, PSNR 28 → 30+
- Densification frozen, opacity and SH refine
- Risk: Overfitting to training views
Phase 4: Final Polish (iter 30,000+)
- Minimal change, diminishing returns
- Risk: Continued training may degrade test views| Failure Mode | Symptom | Root Cause | Fix |
|---|---|---|---|
| Premature plateau | PSNR stops improving by iter 5,000 | Densification frozen too early; lr too low | Increase densify_until_iter; raise lr |
| Never converges | Loss oscillates, PSNR ~20 at iter 30k | Learning rate too high; bad initialization | Reduce lr 10x; check SfM point cloud |
| Train-good/test-bad | High train PSNR, low test PSNR | Overfitting; insufficient camera coverage | More cameras; early stopping; regularization |
| Sudden regression | PSNR drops dramatically mid-training | Gradient explosion; bad ADC clone; data corruption | Check gradient norms; add clipping; verify data |
| Asymmetric convergence | Some views perfect, others terrible | SfM sparse in some regions; uneven camera distribution | Add cameras; increase densification in sparse areas |
| Late-stage degradation | PSNR peaks then declines | Overfitting SH; opacity over-adaptation | Early stopping at peak; reduce opacity_lr |
See references/convergence-trajectories.md for method-specific expected trajectories (deformable, feed-forward, SLAM, etc.).
| Strategy | Description | When to Use | Pitfalls |
|---|---|---|---|
| Data Parallel (DDP) | Each GPU trains full model on different image batch | Standard for large datasets | All-reduce bottleneck with high Gaussian count; requires gradient sync |
| Model Parallel | Split Gaussians across GPUs | When single GPU VRAM insufficient | Load imbalance; complex rasterization coordination |
| Pipeline Parallel | Split training stages across GPUs | Rare for 3DGS | Not well-supported by rasterization kernels |
| FSDP | Shard optimizer states + gradients | Very large Gaussian counts | Overhead for moderate counts; CPU offload needed |
| Bug ID | Symptom | Cause | Fix |
|---|---|---|---|
| DT-01 | Loss diverges on rank 0 only | Gradient sync issue; non-deterministic ADC | Use torch.distributed.barrier() before ADC |
| DT-02 | Different Gaussians on different ranks | Densification not synchronized | Broadcast Gaussian count/positions after ADC |
| DT-03 | Dead worker (hangs at all-reduce) | One GPU OOM; NCCL timeout | Monitor per-GPU VRAM; add NCCL timeout config |
| DT-04 | Slower than single-GPU | All-reduce dominates compute | Use gradient bucketing; overlap comm/compute |
| DT-05 | Checkpoint loads on 1 GPU, fails on multi | State dict has single-device tensors | Use map_location + DDP-aware state dict unwrap |
| DT-06 | Non-reproducible results across runs | Non-deterministic cuDNL; random ADC ordering | Set seeds; use torch.use_deterministic_algorithms(True) |
checkpoint = {
# Model state
'gaussian_params': {
'_xyz': gaussians._xyz.data, # [N, 3]
'_features': gaussians._features.data, # [N, sh_dim]
'_opacity': gaussians._opacity.data, # [N, 1]
'_scaling': gaussians._scaling.data, # [N, 3]
'_rotation': gaussians._rotation.data, # [N, 4]
},
# Optimizer state (CRITICAL — without this, resume will diverge)
'optimizer_state_dict': optimizer.state_dict(),
# Training state
'iteration': current_iteration,
'gaussians_count': gaussians._xyz.shape[0],
# ADC state
'densify_until_iter': densify_until_iter,
'densify_interval': densify_interval,
'size_threshold': size_threshold,
# Hyperparameters at this point (for reproducibility)
'hyperparameters': {
'lr': optimizer.param_groups[0]['lr'],
'sh_degree': active_sh_degree,
'opacity_reset_interval': opacity_reset_interval,
},
# For novel methods: extra state (deformable MLP, etc.)
'extra_state': extra_module.state_dict() if extra_module else None,
}| Issue | Symptom | Cause | Fix |
|---|---|---|---|
| Resume diverges | Loss jumps or NaN after resume | Optimizer state not saved/loaded | Always save + restore optimizer state_dict |
| Wrong Gaussian count | Gaussians mismatch on resume | ADC ran between save and resume | Save AFTER ADC cycle, not during |
| Shape mismatch | Tensor size error on load | Pruning changed Gaussian count | Save count explicitly; handle add/remove |
| SH degree mismatch | Feature dim error | SH degree auto-incremented during training | Save active_sh_degree; restore it |
| Scale/rotation mismatch | Bad rendering after resume | Activation functions applied during save | Save pre-activation values (_scaling, _rotation) |
| Novel method state lost | Deformation MLP reset on resume | Extra module state not in checkpoint | Include all module state_dicts in checkpoint |
| Method Family | Common Stability Issue | Root Cause | Mitigation |
|---|---|---|---|
| Deformable GS (4DGS, Deformable-3DGS) | Deformation MLP outputs NaN | Unconstrained MLP output; large gradients through time | Add tanh activation on output; gradient clip; warmup with frozen base |
| Feed-forward GS (pixelSplat, LRM-based) | Instability with few training images | Model predicts Gaussians from sparse views | More training iterations; 2D feature regularization |
| MoE-GS | Expert collapse (all routing to 1 expert) | Router imbalance; load balancing loss weight too low | Increase load-balancing loss; add router z-loss |
| Physics-based GS (PhysGaussian, Springs) | Physics simulation diverges | Large time step; unstable integrator | Reduce dt; use semi-implicit Euler; add damping |
| SLAM-GS | Drift accumulation over time | Incremental map update without global optimization | Periodic global BA; keyframe-based adjustment |
| Compression GS | Quality collapse after pruning | Pruned critical Gaussians | Importance-aware pruning; fine-tune after prune |
| GaussianGrasper / Embodied | Grasp success drops during training | Sim-to-real gap amplifies | Domain randomization; curriculum learning |
| PBR / Material GS | Material decomposition unstable | Joint optimization of geometry + material under-determined | Stage training: geometry first, then material |
| City-scale / Large-scene | OOM or spatial discontinuity | Too many Gaussians in one scene | Block-wise training; LOD hierarchy |
| GaussTrace / Provenance | Provenance tags mismatch after ADC | Clone/split not propagating tags | Custom ADC that preserves provenance metadata |
torch.autograd.gradcheck or manual gradient norm logging for the novel module.This skill detects 60 runtime failure patterns (as opposed to the code-reviewer's 104 static code bugs). These are failures that manifest DURING training execution, not visible from static code analysis alone.
| Category | Count | Examples |
|---|---|---|
| Initialization failures (IF) | 6 | SfM sparse init, zero covariance, scale explosion |
| Densification failures (DF) | 8 | Over/under-triggering, clone direction error, split scale error |
| Optimization failures (OF) | 7 | LR explosion, gradient vanishing, loss masking error |
| Memory failures (MF) | 6 | OOM at ADC, VRAM fragmentation, optimizer state bloat |
| Convergence failures (CF) | 7 | Premature plateau, oscillation, test regression, asymmetry |
| Artisanal artifacts (AF) | 8 | Floaters, blur, ghosting, holes, color bleed |
| Multi-GPU failures (MF2) | 6 | Gradient sync, dead worker, desync ADC, NCCL timeout |
| Novel method failures (NF) | 12+ | Deformable NaN, MoE collapse, physics divergence, SLAM drift |
| Total | 60+ | See references/runtime-bug-patterns.md |
For the full pattern database with symptoms, root causes, diagnostics, and fixes, see references/runtime-bug-patterns.md.
Before presenting any diagnosis, verify:
If any SC check fails: Do NOT present the diagnosis. Re-examine the failed check and re-run from SC-1.
The following are categorical prohibitions. Violating any of these invalidates the output:
Do NOT try to apply the logic, bug patterns, convergence data, or technical details described in this skill from memory. Always read the SKILL.md and referenced files from disk before producing any output. The knowledge base is updated frequently; stale memory may produce outdated, inaccurate, or fabricated results.
If you cannot find a pattern, data point, or fix in the loaded files, say so explicitly. Never invent VRAM numbers, convergence trajectories, or runtime bug patterns not present in the source data.
If you like it, please star this repo https://github.com/jaccen/Awesome-Gaussian-Skills
© jaccen, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 3 other files (references) in skills/3dgs-training-debugger of jaccen/Awesome-Gaussian-Skills.
Open the folder on GitHubat commit 437c820
3dgs Training Debugger next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| 3dgs Training Debugger this skilljaccen/Awesome-Gaussian-Skills | 161 | — | ~7k | Automated safety check: Pass | Apache-2.0 | |
| AristotleK-Dense-AI/mimeographs | 129 | — | ~1.6k | Automated safety check: Pass | MIT | |
| Codd Fable Consultyohey-w/codd-dev | 115 | — | ~3.8k | Automated safety check: Pass | MIT | |
| Swarm Migrateyonatangross/orchestkit | 289 | — | ~3.7k | Automated safety check: Notes | MIT | |
| Revision WorkflowBingHanOfUESTC/open_agent_team | 106 | — | ~264 | Automated safety check: Pass | MIT | |
| Python Pypi Package Buildergithub/awesome-copilot | 40k | 1 repos | ~4.6k | Automated safety check: Pass | MIT |
K-Dense-AI/mimeographs
Applies the frameworks of Aristotle (ancient Greek philosopher, logic, ethics, metaphysics, 384-322 BCE) to decision-making, ethics, and analysis.
yohey-w/codd-dev
Consult Claude Fable 5 (Anthropic's most capable model for deep, long-horizon reasoning — or whichever model currently holds that role) for a genuinely hard, novel CoDD architectural design…
yonatangross/orchestkit
Cross-repo migration swarm — one coordinator + N parallel subagents (one per target repo) that apply the same transformation, open PRs, wait for CI, and report back to a shared JSON ledger.
BingHanOfUESTC/open_agent_team
Versioned fiction revision workflow. An agent skill from BingHanOfUESTC/open_agent_team.
github/awesome-copilot
End-to-end skill for building, testing, linting, versioning, and publishing a production-grade Python library to PyPI.
SpillwaveSolutions/agent-brain
Modern Python coaching covering language foundations through advanced production patterns.
jaccen/Awesome-Gaussian-Skills
Review 3DGS implementation code for correctness, performance bugs, and best practices.
jaccen/Awesome-Gaussian-Skills
Generate CN patent docs (claims, specification, abstract) and software copyright materials from AI/big-data project code or docs.
jaccen/Awesome-Gaussian-Skills
3DGS Articulated Object Reasoning & Digital Twin Agent. An agent skill from jaccen/Awesome-Gaussian-Skills.
jaccen/Awesome-Gaussian-Skills
3DGS compression-to-deployment pipeline: quantization (scalar/VQ/mixed-precision), pruning (coreset/adaptive/variational/merge/Bayesian), progressive streaming & LoD, Web/WebGPU/mobile deployment…
jaccen/Awesome-Gaussian-Skills
MCP protocol integration with 3DGS rendering pipeline: Agent-controlled Three.js/WebGPU rendering, voice-driven scene reconstruction, real-time parameter manipulation, light tracing backend.
jaccen/Awesome-Gaussian-Skills
Read and summarize 3DGS research papers. An agent skill from jaccen/Awesome-Gaussian-Skills.
Categories
Diagnose and fix 3DGS training-time failures: NaN losses, OOM crashes, divergent optimization, floater artifacts, densification failures, hyperparameter sensitivity. 3dgs Training Debugger is an agent skill from jaccen/Awesome-Gaussian-Skills. Diagnose and fix 3DGS training-time failures: NaN losses, OOM crashes, divergent optimization, floater artifacts, densification failures, hyperparameter sensitivity.
3dgs Training Debugger fits situations like: : 3DGS training crashes; produces poor results; loss is NaN/Inf; VRAM exhaustion.
Run `npx skills add jaccen/Awesome-Gaussian-Skills --skill 3dgs-training-debugger -a claude-code`. Or copy the skill folder (skills/3dgs-training-debugger in jaccen/Awesome-Gaussian-Skills) into .claude/skills/3dgs-training-debugger in your project. Claude Code loads it when a task matches its description.
Run `npx skills add jaccen/Awesome-Gaussian-Skills --skill 3dgs-training-debugger -a codex`. Or copy the skill folder (skills/3dgs-training-debugger in jaccen/Awesome-Gaussian-Skills) into .agents/skills/3dgs-training-debugger in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jaccen/Awesome-Gaussian-Skills --skill 3dgs-training-debugger -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/3dgs-training-debugger, .gemini/skills/3dgs-training-debugger, .github/skills/3dgs-training-debugger and .opencode/skills/3dgs-training-debugger in your project.
SKILL.md names no scripts, command-line tools or credentials: 3dgs Training Debugger is instructions for the agent only.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
3dgs Training Debugger is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 7k tokens (SKILL.md is roughly 28k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 12k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with 3dgs Training Debugger: Aristotle (K-Dense-AI/mimeographs, 129 stars), Codd Fable Consult (yohey-w/codd-dev, 115 stars), Swarm Migrate (yonatangross/orchestkit, 289 stars) and Revision Workflow (BingHanOfUESTC/open_agent_team, 106 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
jaccen (a GitHub user) maintains it in jaccen/Awesome-Gaussian-Skills, which has 161 GitHub stars. The repository holds 13 skills in this directory. The repository was last updated on October 8, 2026.
Source: jaccen/Awesome-Gaussian-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.