Agent skill

GPU Memory Model

by mohitmishra786 in mohitmishra786/low-level-dev-skills

GPU memory model skill for SIMT execution and memory hierarchy.

MITAuto-check passedAgent Workflows

Install GPU Memory Model

skills CLI
$ npx skills add mohitmishra786/low-level-dev-skills --skill gpu-memory-model -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install mohitmishra786/low-level-dev-skills gpu-memory-model --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/mohitmishra786/low-level-dev-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/gpu/gpu-memory-model .claude/skills/gpu-memory-model && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
gpu-memory-model
GitHub stars
253
Token cost
~1.9k tokens
SKILL.md length
560 words
Files
1
Skills in repo
138
Repo updated
First seen
Licence
MIT

At a glance

GPU memory model skill for SIMT execution and memory hierarchy.

  • Works in 9 steps: SIMT execution model → Warp vs wavefront → Warp divergence cost model → …
  • Analyzing warp divergence
  • SKILL.md covers Purpose, When to Use, Workflow and Common Problems, plus 1 more section
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

GPU Memory Model is an agent skill from mohitmishra786/low-level-dev-skills. GPU memory model skill for SIMT execution and memory hierarchy. Use when analyzing warp divergence, memory coalescing, shared memory bank conflicts, cache behavior, atomics, or occupancy tradeoffs. Activates on queries about SIMT, warp coalescing, bank conflicts, wavefront, GPU occupancy, or memory-bound kernels.

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Agent Workflows, covering Agent memory. It works with NVIDIA AI Platform. The repository describes itself as: A curated suite of AI agent skills for systems and low-level programming with C/C++, Rust, and Zig toolchains, covering compilers, debuggers, profilers, build systems…. The licence is MIT.

When your agent uses it

  • Analyzing warp divergence
  • Memory coalescing
  • Shared memory bank conflicts
  • Occupancy tradeoffs

Example prompts

  • “/gpu-memory-model”

Workflow steps

9 steps, taken from the step headings in SKILL.md.

  1. SIMT execution model
  2. Warp vs wavefront
  3. Warp divergence cost model
  4. Global memory coalescing
  5. Shared memory bank conflicts
  6. L1/L2 cache behavior
  7. Atomics and memory ordering
  8. Occupancy vs latency hiding
  9. Quick reference table

What it can do on your machine

Read from SKILL.md and the folder at commit bdc5847. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are c and bash).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

GPU Memory Model loads about 1.9k tokens when it runs. Until then it costs about 83 tokens; SKILL.md has 560 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~83
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from mohitmishra786/low-level-dev-skills at commit bdc5847, republished under its MIT licence (© mohitmishra786). 560 words, ~1,937 tokens.

Download SKILL.mdSave it as .claude/skills/gpu-memory-model/SKILL.md (or your agent's skills folder).
name
gpu-memory-model
description
GPU memory model skill for SIMT execution and memory hierarchy. Use when analyzing warp divergence, memory coalescing, shared memory bank conflicts, cache behavior, atomics, or occupancy tradeoffs. Activates on queries about SIMT, warp coalescing, bank conflicts, wavefront, GPU occupancy, or memory-bound kernels.

GPU Memory Model

Purpose

Explain the GPU execution and memory model for agents optimizing kernels: SIMT execution, warp (32) vs wavefront (64) divergence costs, global memory coalescing rules, shared memory bank conflicts, L1/L2 cache behavior, atomic memory ordering, and the occupancy-vs-latency-hiding tradeoff.

When to Use

  • Diagnosing why a kernel is memory-bound despite high theoretical bandwidth
  • Understanding warp divergence from branching
  • Fixing shared memory bank conflicts in tiled algorithms
  • Choosing block size for occupancy vs register pressure
  • Porting kernels between NVIDIA (warp 32) and AMD (wavefront 64)
  • Reasoning about atomic contention in parallel reductions

Workflow

1. SIMT execution model
GPU hardware
├── Device
│   └── SM / CU (Streaming Multiprocessor / Compute Unit)
│       ├── Warp schedulers (NVIDIA) or Wavefront schedulers (AMD)
│       │   └── Warp/Wavefront (32 or 64 threads in lockstep)
│       ├── Register file (partitioned per thread)
│       ├── Shared memory / LDS (per SM)
│       └── L1 cache (often shared with shared memory)
└── L2 cache (device-wide) → DRAM/HBM

SIMT (Single Instruction, Multiple Threads): one instruction stream drives a warp/wavefront; each thread has its own registers and thread ID but executes the same instruction in lockstep.

2. Warp vs wavefront
VendorUnit sizeName
NVIDIA32 threadsWarp
AMD64 threadsWavefront

Implications:

  • Reduction trees: NVIDIA halves at 16→8→4→2→1; AMD at 32→16→8→4→2→1
  • Block sizes: prefer multiples of 32 (NVIDIA) or 64 (AMD)
  • Occupancy counters report active warps/wavefronts per SM
3. Warp divergence cost model

When threads in a warp take different branches, the hardware serializes paths:

c
// Divergent: half warp does A, half does B → 2x instruction issue
if (threadIdx.x % 2 == 0) {
    result = expensive_a(data[idx]);
} else {
    result = expensive_b(data[idx]);
}

// Non-divergent: all threads same path
result = expensive_a(data[idx]);

Mitigations:

  • Predication: compute both, select with ?: (trade compute for uniformity)
  • Separate kernels for different code paths
  • Branch only on block-level data (uniform within warp)
  • Loop over bins instead of if (data[i] < threshold) per thread with scattered outcomes

Divergence cost ≈ sum of paths taken (not max).

4. Global memory coalescing

NVIDIA coalescing rule (simplified): threads in a warp accessing consecutive 4-byte words → single 128-byte transaction.

c
// Coalesced: consecutive threads → consecutive addresses
int idx = blockIdx.x * blockDim.x + threadIdx.x;
float val = data[idx];

// Uncoalesced: stride access
float val = data[threadIdx.x * stride];  // stride > 1

// Partially coalesced: misaligned start
float val = data[base + threadIdx.x * 3];

AoS vs SoA impact:

c
// AoS — poor coalescing when reading one field
struct Particle { float x, y, z; };
float x = particles[i].x;  // threads read with stride 3

// SoA — coalesced
float x = pos_x[i];
5. Shared memory bank conflicts

Shared memory is divided into 32 banks (4-byte words). Simultaneous accesses to different addresses in the same bank serialize.

c
__shared__ float tile[32][32];

// Bank conflict: all threads access tile[threadIdx.x][0]
// 32 threads, 32 banks, but column 0 → same bank per row offset
float val = tile[threadIdx.x][0];

// Fix: pad columns to break bank alignment
__shared__ float tile[32][33];  // +1 padding

Detection: Nsight Compute l1tex__data_bank_conflicts_pipe_lsu_mem_shared_op_ld.sum or NCU shared load conflict metrics.

6. L1/L2 cache behavior
LevelScopeNotes
L1Per-SMOften unified with shared mem; configurable split
L2Device-wideCache lines typically 128 bytes
Texture/L1 readonlyPer-SMCached read-only path for uniform access

Cache-friendly patterns:

  • Spatial locality: consecutive threads access consecutive memory
  • Temporal locality: reuse data in shared memory before re-fetching global
  • Avoid random scatter: atomic updates and pointer chasing defeat caches
c
// Cache-friendly tile load
for (int t = 0; t < num_tiles; t++) {
    __shared__ float smem[TILE][TILE];
    smem[ty][tx] = global[row * N + t * TILE + tx];
    __syncthreads();
    // compute from smem — L1/L2 only hit on first load per tile
}
Show full SKILL.md (212 more words)Show less
7. Atomics and memory ordering

GPU atomics (atomicAdd, atomicCAS, atomicExch) provide sequential consistency among threads targeting the same address, but high contention serializes execution.

c
// Bad: all threads atomic to one counter
atomicAdd(&global_sum, local_val);

// Better: per-block reduction, one atomic per block
__shared__ float block_sum;
// ... warp reduce to block_sum ...
if (threadIdx.x == 0)
    atomicAdd(&global_sum, block_sum);

HIP/CUDA memory fences:

c
__threadfence_block();  // visible to threads in same block
__threadfence();        // visible to all threads on device
__threadfence_system(); // visible to host (expensive)
8. Occupancy vs latency hiding
Occupancy tradeoff
├── High occupancy → more warps to hide memory latency
│   └── Costs: fewer registers/SM, less shared mem per block
└── Low occupancy + high ILP → enough independent instructions per warp
    └── Works for compute-bound kernels with deep pipelines

Decision tree:

Memory-bound kernel?
├── Yes → maximize active warps (occupancy), coalesce, tile with shared mem
└── No (compute-bound) → may lower occupancy if registers enable more ILP
bash
# Measure achieved occupancy
ncu --metrics sm__warps_active.avg.pct_of_peak_sustained_active ./app

Rule of thumb: memory-bound kernels need occupancy ≥ 50%; compute-bound may run well at 25% with sufficient instruction-level parallelism.

9. Quick reference table
SymptomLikely causeCheck
Low DRAM throughputUncoalesced accessNCU memory workload analysis
Shared load stallsBank conflictsPad arrays; change access pattern
High stall: barrierMissing __syncthreads or divergence at barriersynccheck
Atomic bottleneckToo many contended atomicsHierarchical reduction
Low occupancyRegister/shared mem pressurelaunch__occupancy_limit_* metrics

Common Problems

SymptomCauseFix
10x slower after SoA→AoS changeStrided coalescing brokenKeep hot fields in SoA layout
Tiled matmul slower than naiveBank conflicts in shared tilePad shared array columns
Identical code, different perf NVIDIA vs AMDWarp 32 vs wavefront 64Retune block size and reductions
Atomics correct but slowGlobal contentionBlock-level reduce first
High L2 hit but still slowL2 bandwidth saturatedReduce total bytes moved
  • skills/gpu/cuda — practical kernel patterns using this model
  • skills/gpu/cuda-profiling — metrics to validate coalescing and occupancy
  • skills/gpu/hip-rocm — AMD wavefront-specific tuning
  • skills/low-level-programming/cpu-cache-opt — CPU cache concepts (analogous)
  • skills/low-level-programming/simd-intrinsics — vector width on CPU side
  • skills/profilers/hardware-counters — general cache miss measurement

© mohitmishra786, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/gpu/gpu-memory-model of mohitmishra786/low-level-dev-skills.

Open the folder on GitHubat commit bdc5847

Compare with similar skills

GPU Memory Model next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

GPU Memory Model compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
GPU Memory Model this skillmohitmishra786/low-level-dev-skills253—~1.9kAutomated safety check: PassMIT
Project Timeline Reportthedotmack/claude-mem97k1 repos~3.1kAutomated safety check: PassApache-2.0
Neat-Freak Knowledge CloseoutKKKKhazix/khazix-skills21k—~1.9kAutomated safety check: PassMIT
Beads Task Memorygastownhall/beads28k—~1.2kAutomated safety check: PassMIT
MemPalace Memory SearchMemPalace/mempalace59k—~1.4kAutomated safety check: PassMIT
Reflect on Session Learningscursor/plugins10k5 repos~1.2kAutomated safety check: PassNone

Similar skills

  • Project Timeline Report

    thedotmack/claude-mem

    Writes a narrative Journey Into report on a project's whole development history, built from the timeline that claude-mem has recorded.

    97k GitHub starsUsed in 1 repo~3.1k tokens
    Agent WorkflowsAuto-check passed
  • Neat-Freak Knowledge Closeout

    KKKKhazix/khazix-skills

    Brings project docs, agent rule files, authorized memory and leftover workspace files back in line with what the code and runtime actually do at the end of a work session.

    21k GitHub stars~1.9k tokensUpdated 6 days ago
    Agent WorkflowsAuto-check passed
  • Beads Task Memory

    gastownhall/beads

    Tracks multi-session work with dependencies in the bd issue tracker so the agent can find ready tasks and recover its context after conversation compaction.

    28k GitHub stars~1.2k tokensUpdated today
    Agent WorkflowsAuto-check passed
  • MemPalace Memory Search

    MemPalace/mempalace

    Mines project files and conversation exports into a local, searchable memory palace and recalls past work by semantic search through the mempalace CLI.

    59k GitHub stars~1.4k tokensUpdated today
    Agent WorkflowsAuto-check passed
  • Official

    Starts three parallel reviewer subagents over the current conversation transcript, then turns their findings into concrete edits to existing skills.

    10k GitHub starsUsed in 5 repos~1.2k tokens
    Agent WorkflowsAuto-check passed
  • Compound Learning Writer

    EveryInc/compound-engineering-plugin

    Records one solved and verified problem as a durable learning in the repository, but only when the reasoning is not already clear from the final code, tests or docs.

    25k GitHub stars~2k tokensUpdated today
    Agent WorkflowsAuto-check passed

More from mohitmishra786/low-level-dev-skills

All 138 skills in this repo
  • ARM and AArch64 Assembly

    mohitmishra786/low-level-dev-skills

    Guides reading and writing AArch64 and ARM Thumb assembly: compiler output, inline asm, registers, the AAPCS calling convention and NEON or SVE basics.

    253 GitHub stars~1.9k tokensUpdated 3 mo ago
    Auto-check passed
  • RISC-V Assembly Guide

    mohitmishra786/low-level-dev-skills

    Reference for RISC-V assembly on RV32 and RV64: register names and calling convention, extension naming, GCC and Clang inline asm, and QEMU with GDB debugging.

    253 GitHub stars~1.8k tokensUpdated 3 mo ago
    Auto-check passed
  • x86-64 Assembly Reference

    mohitmishra786/low-level-dev-skills

    Explains x86-64 registers, the System V AMD64 calling convention, and how to read compiler-generated or inline assembly.

    253 GitHub stars~1.5k tokensUpdated 3 mo ago
    Auto-check passed
  • Bazel for C and C++

    mohitmishra786/low-level-dev-skills

    Guides your agent through Bazel for C/C++ projects: BUILD files, Bzlmod dependencies, toolchain registration, remote execution, dependency queries and sandbox debugging.

    253 GitHub stars~1.5k tokensUpdated 3 mo ago
    Auto-check passed
  • Binary Hardening

    mohitmishra786/low-level-dev-skills

    Binary hardening skill for security-hardened C/C++ builds. An agent skill from mohitmishra786/low-level-dev-skills.

    253 GitHub stars~2k tokensUpdated 3 mo ago
    Auto-check passed
  • Binutils

    mohitmishra786/low-level-dev-skills

    GNU binutils skill for binary manipulation and analysis. An agent skill from mohitmishra786/low-level-dev-skills.

    253 GitHub stars~1.2k tokensUpdated 3 mo ago
    Auto-check passed

Categories

Questions about GPU Memory Model

What does GPU Memory Model do?

GPU memory model skill for SIMT execution and memory hierarchy. GPU Memory Model is an agent skill from mohitmishra786/low-level-dev-skills. GPU memory model skill for SIMT execution and memory hierarchy.

When should I use GPU Memory Model?

GPU Memory Model fits situations like: analyzing warp divergence; memory coalescing; shared memory bank conflicts; occupancy tradeoffs.

How do I install GPU Memory Model in Claude Code?

Run `npx skills add mohitmishra786/low-level-dev-skills --skill gpu-memory-model -a claude-code`. Or copy the skill folder (skills/gpu/gpu-memory-model in mohitmishra786/low-level-dev-skills) into .claude/skills/gpu-memory-model in your project. Claude Code loads it when a task matches its description.

How do I install GPU Memory Model in Codex?

Run `npx skills add mohitmishra786/low-level-dev-skills --skill gpu-memory-model -a codex`. Or copy the skill folder (skills/gpu/gpu-memory-model in mohitmishra786/low-level-dev-skills) into .agents/skills/gpu-memory-model in your project. Codex loads it when a task matches its description.

Can I use GPU Memory Model in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mohitmishra786/low-level-dev-skills --skill gpu-memory-model -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/gpu-memory-model, .gemini/skills/gpu-memory-model, .github/skills/gpu-memory-model and .opencode/skills/gpu-memory-model in your project.

What does GPU Memory Model need to run?

SKILL.md names no scripts, command-line tools or credentials: GPU Memory Model is instructions for the agent only.

Does GPU Memory Model access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is GPU Memory Model safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does GPU Memory Model use?

GPU Memory Model is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does GPU Memory Model use?

About 1.9k tokens (SKILL.md is roughly 7.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to GPU Memory Model?

Skills that share tags, products or a category with GPU Memory Model: Project Timeline Report (thedotmack/claude-mem, 97k stars), Neat-Freak Knowledge Closeout (KKKKhazix/khazix-skills, 21k stars), Beads Task Memory (gastownhall/beads, 28k stars) and MemPalace Memory Search (MemPalace/mempalace, 59k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains GPU Memory Model?

mohitmishra786 (a GitHub user) maintains it in mohitmishra786/low-level-dev-skills, which has 253 GitHub stars. The repository holds 138 skills in this directory. The repository was last updated on June 27, 2026.

Source: mohitmishra786/low-level-dev-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.