Project Timeline Report
thedotmack/claude-mem
Writes a narrative Journey Into report on a project's whole development history, built from the timeline that claude-mem has recorded.
GPU memory model skill for SIMT execution and memory hierarchy.
$ npx skills add mohitmishra786/low-level-dev-skills --skill gpu-memory-model -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install mohitmishra786/low-level-dev-skills gpu-memory-model --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/mohitmishra786/low-level-dev-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/gpu/gpu-memory-model .claude/skills/gpu-memory-model && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "gpu-memory-model" agent skill from https://github.com/mohitmishra786/low-level-dev-skills/tree/main/skills/gpu/gpu-memory-model into .claude/skills/gpu-memory-model/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gpu-memory-model", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/mohitmishra786/low-level-dev-skills/tree/main/skills/gpu/gpu-memory-modelType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add mohitmishra786/low-level-dev-skills --skill gpu-memory-model -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install mohitmishra786/low-level-dev-skills gpu-memory-model --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mohitmishra786/low-level-dev-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/gpu/gpu-memory-model .agents/skills/gpu-memory-model && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "gpu-memory-model" agent skill from https://github.com/mohitmishra786/low-level-dev-skills/tree/main/skills/gpu/gpu-memory-model into .agents/skills/gpu-memory-model/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gpu-memory-model", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add mohitmishra786/low-level-dev-skills --skill gpu-memory-model -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install mohitmishra786/low-level-dev-skills gpu-memory-model --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mohitmishra786/low-level-dev-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/gpu/gpu-memory-model .cursor/skills/gpu-memory-model && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "gpu-memory-model" agent skill from https://github.com/mohitmishra786/low-level-dev-skills/tree/main/skills/gpu/gpu-memory-model into .cursor/skills/gpu-memory-model/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gpu-memory-model", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/mohitmishra786/low-level-dev-skills.git --path skills/gpu/gpu-memory-model--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add mohitmishra786/low-level-dev-skills --skill gpu-memory-model -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install mohitmishra786/low-level-dev-skills gpu-memory-model --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mohitmishra786/low-level-dev-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/gpu/gpu-memory-model .gemini/skills/gpu-memory-model && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "gpu-memory-model" agent skill from https://github.com/mohitmishra786/low-level-dev-skills/tree/main/skills/gpu/gpu-memory-model into .gemini/skills/gpu-memory-model/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gpu-memory-model", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install mohitmishra786/low-level-dev-skills gpu-memory-modelInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add mohitmishra786/low-level-dev-skills --skill gpu-memory-model -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/mohitmishra786/low-level-dev-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/gpu/gpu-memory-model .github/skills/gpu-memory-model && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "gpu-memory-model" agent skill from https://github.com/mohitmishra786/low-level-dev-skills/tree/main/skills/gpu/gpu-memory-model into .github/skills/gpu-memory-model/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gpu-memory-model", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add mohitmishra786/low-level-dev-skills --skill gpu-memory-model -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install mohitmishra786/low-level-dev-skills gpu-memory-model --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mohitmishra786/low-level-dev-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/gpu/gpu-memory-model .opencode/skills/gpu-memory-model && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "gpu-memory-model" agent skill from https://github.com/mohitmishra786/low-level-dev-skills/tree/main/skills/gpu/gpu-memory-model into .opencode/skills/gpu-memory-model/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gpu-memory-model", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
gpu-memory-modelGPU memory model skill for SIMT execution and memory hierarchy.
GPU Memory Model is an agent skill from mohitmishra786/low-level-dev-skills. GPU memory model skill for SIMT execution and memory hierarchy. Use when analyzing warp divergence, memory coalescing, shared memory bank conflicts, cache behavior, atomics, or occupancy tradeoffs. Activates on queries about SIMT, warp coalescing, bank conflicts, wavefront, GPU occupancy, or memory-bound kernels.
Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Agent Workflows, covering Agent memory. It works with NVIDIA AI Platform. The repository describes itself as: A curated suite of AI agent skills for systems and low-level programming with C/C++, Rust, and Zig toolchains, covering compilers, debuggers, profilers, build systems…. The licence is MIT.
9 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit bdc5847. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are c and bash).
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
GPU Memory Model loads about 1.9k tokens when it runs. Until then it costs about 83 tokens; SKILL.md has 560 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from mohitmishra786/low-level-dev-skills at commit bdc5847, republished under its MIT licence (© mohitmishra786). 560 words, ~1,937 tokens.
.claude/skills/gpu-memory-model/SKILL.md (or your agent's skills folder).Explain the GPU execution and memory model for agents optimizing kernels: SIMT execution, warp (32) vs wavefront (64) divergence costs, global memory coalescing rules, shared memory bank conflicts, L1/L2 cache behavior, atomic memory ordering, and the occupancy-vs-latency-hiding tradeoff.
GPU hardware
├── Device
│ └── SM / CU (Streaming Multiprocessor / Compute Unit)
│ ├── Warp schedulers (NVIDIA) or Wavefront schedulers (AMD)
│ │ └── Warp/Wavefront (32 or 64 threads in lockstep)
│ ├── Register file (partitioned per thread)
│ ├── Shared memory / LDS (per SM)
│ └── L1 cache (often shared with shared memory)
└── L2 cache (device-wide) → DRAM/HBMSIMT (Single Instruction, Multiple Threads): one instruction stream drives a warp/wavefront; each thread has its own registers and thread ID but executes the same instruction in lockstep.
| Vendor | Unit size | Name |
|---|---|---|
| NVIDIA | 32 threads | Warp |
| AMD | 64 threads | Wavefront |
Implications:
When threads in a warp take different branches, the hardware serializes paths:
// Divergent: half warp does A, half does B → 2x instruction issue
if (threadIdx.x % 2 == 0) {
result = expensive_a(data[idx]);
} else {
result = expensive_b(data[idx]);
}
// Non-divergent: all threads same path
result = expensive_a(data[idx]);Mitigations:
?: (trade compute for uniformity)if (data[i] < threshold) per thread with scattered outcomesDivergence cost ≈ sum of paths taken (not max).
NVIDIA coalescing rule (simplified): threads in a warp accessing consecutive 4-byte words → single 128-byte transaction.
// Coalesced: consecutive threads → consecutive addresses
int idx = blockIdx.x * blockDim.x + threadIdx.x;
float val = data[idx];
// Uncoalesced: stride access
float val = data[threadIdx.x * stride]; // stride > 1
// Partially coalesced: misaligned start
float val = data[base + threadIdx.x * 3];AoS vs SoA impact:
// AoS — poor coalescing when reading one field
struct Particle { float x, y, z; };
float x = particles[i].x; // threads read with stride 3
// SoA — coalesced
float x = pos_x[i];Shared memory is divided into 32 banks (4-byte words). Simultaneous accesses to different addresses in the same bank serialize.
__shared__ float tile[32][32];
// Bank conflict: all threads access tile[threadIdx.x][0]
// 32 threads, 32 banks, but column 0 → same bank per row offset
float val = tile[threadIdx.x][0];
// Fix: pad columns to break bank alignment
__shared__ float tile[32][33]; // +1 paddingDetection: Nsight Compute l1tex__data_bank_conflicts_pipe_lsu_mem_shared_op_ld.sum or NCU shared load conflict metrics.
| Level | Scope | Notes |
|---|---|---|
| L1 | Per-SM | Often unified with shared mem; configurable split |
| L2 | Device-wide | Cache lines typically 128 bytes |
| Texture/L1 readonly | Per-SM | Cached read-only path for uniform access |
Cache-friendly patterns:
// Cache-friendly tile load
for (int t = 0; t < num_tiles; t++) {
__shared__ float smem[TILE][TILE];
smem[ty][tx] = global[row * N + t * TILE + tx];
__syncthreads();
// compute from smem — L1/L2 only hit on first load per tile
}GPU atomics (atomicAdd, atomicCAS, atomicExch) provide sequential consistency among threads targeting the same address, but high contention serializes execution.
// Bad: all threads atomic to one counter
atomicAdd(&global_sum, local_val);
// Better: per-block reduction, one atomic per block
__shared__ float block_sum;
// ... warp reduce to block_sum ...
if (threadIdx.x == 0)
atomicAdd(&global_sum, block_sum);HIP/CUDA memory fences:
__threadfence_block(); // visible to threads in same block
__threadfence(); // visible to all threads on device
__threadfence_system(); // visible to host (expensive)Occupancy tradeoff
├── High occupancy → more warps to hide memory latency
│ └── Costs: fewer registers/SM, less shared mem per block
└── Low occupancy + high ILP → enough independent instructions per warp
└── Works for compute-bound kernels with deep pipelinesDecision tree:
Memory-bound kernel?
├── Yes → maximize active warps (occupancy), coalesce, tile with shared mem
└── No (compute-bound) → may lower occupancy if registers enable more ILP# Measure achieved occupancy
ncu --metrics sm__warps_active.avg.pct_of_peak_sustained_active ./appRule of thumb: memory-bound kernels need occupancy ≥ 50%; compute-bound may run well at 25% with sufficient instruction-level parallelism.
| Symptom | Likely cause | Check |
|---|---|---|
| Low DRAM throughput | Uncoalesced access | NCU memory workload analysis |
| Shared load stalls | Bank conflicts | Pad arrays; change access pattern |
| High stall: barrier | Missing __syncthreads or divergence at barrier | synccheck |
| Atomic bottleneck | Too many contended atomics | Hierarchical reduction |
| Low occupancy | Register/shared mem pressure | launch__occupancy_limit_* metrics |
| Symptom | Cause | Fix |
|---|---|---|
| 10x slower after SoA→AoS change | Strided coalescing broken | Keep hot fields in SoA layout |
| Tiled matmul slower than naive | Bank conflicts in shared tile | Pad shared array columns |
| Identical code, different perf NVIDIA vs AMD | Warp 32 vs wavefront 64 | Retune block size and reductions |
| Atomics correct but slow | Global contention | Block-level reduce first |
| High L2 hit but still slow | L2 bandwidth saturated | Reduce total bytes moved |
skills/gpu/cuda — practical kernel patterns using this modelskills/gpu/cuda-profiling — metrics to validate coalescing and occupancyskills/gpu/hip-rocm — AMD wavefront-specific tuningskills/low-level-programming/cpu-cache-opt — CPU cache concepts (analogous)skills/low-level-programming/simd-intrinsics — vector width on CPU sideskills/profilers/hardware-counters — general cache miss measurement© mohitmishra786, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/gpu/gpu-memory-model of mohitmishra786/low-level-dev-skills.
Open the folder on GitHubat commit bdc5847
GPU Memory Model next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| GPU Memory Model this skillmohitmishra786/low-level-dev-skills | 253 | — | ~1.9k | Automated safety check: Pass | MIT | |
| Project Timeline Reportthedotmack/claude-mem | 97k | 1 repos | ~3.1k | Automated safety check: Pass | Apache-2.0 | |
| Neat-Freak Knowledge CloseoutKKKKhazix/khazix-skills | 21k | — | ~1.9k | Automated safety check: Pass | MIT | |
| Beads Task Memorygastownhall/beads | 28k | — | ~1.2k | Automated safety check: Pass | MIT | |
| MemPalace Memory SearchMemPalace/mempalace | 59k | — | ~1.4k | Automated safety check: Pass | MIT | |
| Reflect on Session Learningscursor/plugins | 10k | 5 repos | ~1.2k | Automated safety check: Pass | None |
thedotmack/claude-mem
Writes a narrative Journey Into report on a project's whole development history, built from the timeline that claude-mem has recorded.
KKKKhazix/khazix-skills
Brings project docs, agent rule files, authorized memory and leftover workspace files back in line with what the code and runtime actually do at the end of a work session.
gastownhall/beads
Tracks multi-session work with dependencies in the bd issue tracker so the agent can find ready tasks and recover its context after conversation compaction.
MemPalace/mempalace
Mines project files and conversation exports into a local, searchable memory palace and recalls past work by semantic search through the mempalace CLI.
cursor/plugins
Starts three parallel reviewer subagents over the current conversation transcript, then turns their findings into concrete edits to existing skills.
EveryInc/compound-engineering-plugin
Records one solved and verified problem as a durable learning in the repository, but only when the reasoning is not already clear from the final code, tests or docs.
mohitmishra786/low-level-dev-skills
Guides reading and writing AArch64 and ARM Thumb assembly: compiler output, inline asm, registers, the AAPCS calling convention and NEON or SVE basics.
mohitmishra786/low-level-dev-skills
Reference for RISC-V assembly on RV32 and RV64: register names and calling convention, extension naming, GCC and Clang inline asm, and QEMU with GDB debugging.
mohitmishra786/low-level-dev-skills
Explains x86-64 registers, the System V AMD64 calling convention, and how to read compiler-generated or inline assembly.
mohitmishra786/low-level-dev-skills
Guides your agent through Bazel for C/C++ projects: BUILD files, Bzlmod dependencies, toolchain registration, remote execution, dependency queries and sandbox debugging.
mohitmishra786/low-level-dev-skills
Binary hardening skill for security-hardened C/C++ builds. An agent skill from mohitmishra786/low-level-dev-skills.
mohitmishra786/low-level-dev-skills
GNU binutils skill for binary manipulation and analysis. An agent skill from mohitmishra786/low-level-dev-skills.
Works with
Categories
GPU memory model skill for SIMT execution and memory hierarchy. GPU Memory Model is an agent skill from mohitmishra786/low-level-dev-skills. GPU memory model skill for SIMT execution and memory hierarchy.
GPU Memory Model fits situations like: analyzing warp divergence; memory coalescing; shared memory bank conflicts; occupancy tradeoffs.
Run `npx skills add mohitmishra786/low-level-dev-skills --skill gpu-memory-model -a claude-code`. Or copy the skill folder (skills/gpu/gpu-memory-model in mohitmishra786/low-level-dev-skills) into .claude/skills/gpu-memory-model in your project. Claude Code loads it when a task matches its description.
Run `npx skills add mohitmishra786/low-level-dev-skills --skill gpu-memory-model -a codex`. Or copy the skill folder (skills/gpu/gpu-memory-model in mohitmishra786/low-level-dev-skills) into .agents/skills/gpu-memory-model in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mohitmishra786/low-level-dev-skills --skill gpu-memory-model -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/gpu-memory-model, .gemini/skills/gpu-memory-model, .github/skills/gpu-memory-model and .opencode/skills/gpu-memory-model in your project.
SKILL.md names no scripts, command-line tools or credentials: GPU Memory Model is instructions for the agent only.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
GPU Memory Model is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.9k tokens (SKILL.md is roughly 7.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with GPU Memory Model: Project Timeline Report (thedotmack/claude-mem, 97k stars), Neat-Freak Knowledge Closeout (KKKKhazix/khazix-skills, 21k stars), Beads Task Memory (gastownhall/beads, 28k stars) and MemPalace Memory Search (MemPalace/mempalace, 59k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
mohitmishra786 (a GitHub user) maintains it in mohitmishra786/low-level-dev-skills, which has 253 GitHub stars. The repository holds 138 skills in this directory. The repository was last updated on June 27, 2026.
Source: mohitmishra786/low-level-dev-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.