CPU cache optimization skill for C/C++ and Rust. An agent skill from mohitmishra786/low-level-dev-skills.

MITAuto-check passedDevelopment

Install Cpu Cache Opt

skills CLI
$ npx skills add mohitmishra786/low-level-dev-skills --skill cpu-cache-opt -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install mohitmishra786/low-level-dev-skills cpu-cache-opt --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/mohitmishra786/low-level-dev-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/low-level-programming/cpu-cache-opt .claude/skills/cpu-cache-opt && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
cpu-cache-opt
GitHub stars
253
Token cost
~1.7k tokens
SKILL.md length
243 words
Files
2 (incl. references)
Skills in repo
138
Repo updated
First seen
Licence
MIT

At a glance

CPU cache optimization skill for C/C++ and Rust. An agent skill from mohitmishra786/low-level-dev-skills.

  • Works in 7 steps: Measure cache performance → Cache line basics → AoS vs SoA data layout → …
  • Diagnosing cache misses
  • SKILL.md covers Purpose, Triggers, Workflow and Related skills
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Cpu Cache Opt is an agent skill from mohitmishra786/low-level-dev-skills. CPU cache optimization skill for C/C++ and Rust. Use when diagnosing cache misses, improving data layout for cache efficiency, using perf stat cache counters, understanding false sharing, prefetching, or structuring AoS vs SoA data layouts. Activates on queries about cache misses, cache lines, false sharing, perf cache counters, data layout optimization, prefetch, AoS vs SoA, or L1/L2/L3 cache performance.

Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/cache-counters.md`).

It sits in Development. It works with C++ and Rust. The repository describes itself as: A curated suite of AI agent skills for systems and low-level programming with C/C++, Rust, and Zig toolchains, covering compilers, debuggers, profilers, build systems…. The licence is MIT.

When your agent uses it

  • Diagnosing cache misses
  • Improving data layout for cache efficiency
  • Using perf stat cache counters
  • Understanding false sharing

Example prompts

  • “/cpu-cache-opt”

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Measure cache performance
  2. Cache line basics
  3. AoS vs SoA data layout
  4. Common cache-unfriendly patterns
  5. False sharing
  6. Prefetching
  7. Cache-friendly algorithm design

What it can do on your machine

Read from SKILL.md and the folder at commit bdc5847. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are c and bash).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Cpu Cache Opt loads about 1.7k tokens when it runs, and up to ~2.8k if it reads all its reference files. Until then it costs about 106 tokens; SKILL.md has 243 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~106
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~2.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from mohitmishra786/low-level-dev-skills at commit bdc5847, republished under its MIT licence (© mohitmishra786). 243 words, ~1,686 tokens.

Download SKILL.mdSave it as .claude/skills/cpu-cache-opt/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
cpu-cache-opt
description
CPU cache optimization skill for C/C++ and Rust. Use when diagnosing cache misses, improving data layout for cache efficiency, using perf stat cache counters, understanding false sharing, prefetching, or structuring AoS vs SoA data layouts. Activates on queries about cache misses, cache lines, false sharing, perf cache counters, data layout optimization, prefetch, AoS vs SoA, or L1/L2/L3 cache performance.

CPU Cache Optimization

Purpose

Guide agents through cache-aware programming: diagnosing cache misses with perf, data layout transformations (AoS→SoA), false sharing detection and fixes, prefetching, and cache-friendly algorithm design.

Triggers

  • "My program has high cache miss rates — how do I fix it?"
  • "What is false sharing and how do I detect it?"
  • "Should I use AoS or SoA data layout?"
  • "How do I measure cache performance with perf?"
  • "How do I use __builtin_prefetch?"
  • "My multithreaded program is slower than single-threaded due to cache"

Workflow

1. Measure cache performance
bash
# Basic cache counters
perf stat -e cache-references,cache-misses,cycles,instructions ./prog

# L1/L2/L3 miss breakdown
perf stat -e \
    L1-dcache-load-misses,\
    L1-dcache-loads,\
    L2-dcache-load-misses,\
    LLC-load-misses,\
    LLC-loads \
    ./prog

# Cache miss rate = L1-dcache-load-misses / L1-dcache-loads
# > 5% is concerning; > 20% is severe

# False sharing detection
perf stat -e \
    machine_clears.memory_ordering,\
    mem_load_l3_hit_retired.xsnp_hitm \
    ./prog
2. Cache line basics
  • Cache line size: 64 bytes on x86-64, ARM (most platforms)
  • L1 cache: 32–64 KB, ~4 cycles latency
  • L2 cache: 256 KB–1 MB, ~12 cycles latency
  • L3 cache: 6–64 MB, ~40 cycles latency
  • Main memory: ~200–300 cycles latency
c
// Check cache line size
long cache_line = sysconf(_SC_LEVEL1_DCACHE_LINESIZE);

// Align data to cache line
struct alignas(64) HotData {
    int counter;
    // ... 60 bytes of data that fit in one line
};

// C
typedef struct {
    int x;
} __attribute__((aligned(64))) AlignedData;
3. AoS vs SoA data layout
c
// AoS (Array of Structures) — default layout
struct Particle {
    float x, y, z;     // position (12 bytes)
    float vx, vy, vz;  // velocity (12 bytes)
    float mass;         // (4 bytes)
    int   flags;        // (4 bytes)
};
Particle particles[N];  // Bad for loops that only need position

// Problem: accessing particles[i].x loads x,y,z,vx,vy,vz,mass,flags
// But we only need x,y,z → 75% of loaded data is wasted

// SoA (Structure of Arrays) — cache-friendly for SIMD + sequential access
struct ParticlesSoA {
    float *x, *y, *z;
    float *vx, *vy, *vz;
    float *mass;
    int   *flags;
};

// Accessing x[i] for i=0..N loads 16 consecutive x values → 0% waste
// Also auto-vectorizes better
4. Common cache-unfriendly patterns
c
// BAD: random access (linked list traversal)
Node *node = head;
while (node) {
    process(node->data);
    node = node->next;  // pointer chasing = cache miss per node
}

// BETTER: pool allocate nodes contiguously
// Or: rewrite as contiguous array with indices

// BAD: stride > cache line in matrix traversal
for (int i = 0; i < N; i++)
    for (int j = 0; j < M; j++)
        sum += matrix[j][i];  // column-major access on row-major array

// GOOD: row-major access
for (int i = 0; i < N; i++)
    for (int j = 0; j < M; j++)
        sum += matrix[i][j];

// BAD: large struct with hot + cold fields
struct Record {
    int id;           // hot: accessed every iteration
    char name[128];   // cold: accessed rarely
    int value;        // hot
    char desc[256];   // cold
};

// GOOD: separate hot and cold data
struct RecordHot { int id; int value; };
struct RecordCold { char name[128]; char desc[256]; };
RecordHot hot_data[N];
RecordCold cold_data[N];
5. False sharing

False sharing occurs when two threads write to different variables that share a cache line, causing constant cache-line invalidations.

c
// BAD: counters likely on same cache line (8 bytes each, line = 64 bytes)
int counter_a;  // thread A's counter
int counter_b;  // thread B's counter

// Both on the same cache line → every write invalidates the other thread's cache

// GOOD: pad to separate cache lines
struct alignas(64) PaddedCounter {
    int value;
    char padding[60];  // Ensure next counter is on different cache line
};

PaddedCounter counters[NUM_THREADS];
// Thread i: counters[i].value++

// C++ standard approach
struct alignas(std::hardware_destructive_interference_size) PaddedCounter {
    int value;
};
6. Prefetching

Manual prefetch hints to hide memory latency:

c
#include <immintrin.h>  // or <xmmintrin.h>

// Prefetch for read (locality 0=non-temporal, 3=high temporal)
__builtin_prefetch(ptr, 0, 3);  // prefetch for read, high locality
__builtin_prefetch(ptr, 1, 3);  // prefetch for write, high locality

// SSE prefetch (x86)
_mm_prefetch((char*)ptr, _MM_HINT_T0);   // L1
_mm_prefetch((char*)ptr, _MM_HINT_T1);   // L2
_mm_prefetch((char*)ptr, _MM_HINT_T2);   // L3
_mm_prefetch((char*)ptr, _MM_HINT_NTA);  // non-temporal (streaming)

// Typical pattern: prefetch N iterations ahead
#define PREFETCH_DIST 8
for (int i = 0; i < N; i++) {
    if (i + PREFETCH_DIST < N)
        __builtin_prefetch(&data[i + PREFETCH_DIST], 0, 3);
    process(data[i]);
}

Prefetching rules:

  • Prefetch too early = cache evicted before use
  • Prefetch too late = no benefit
  • Prefetch distance = memory latency / time per iteration (typically 8–32 elements)
7. Cache-friendly algorithm design
c
// Loop blocking / tiling for matrix operations
// Process cache-fitting blocks instead of full rows/columns
#define BLOCK 64  // tuned to L1 cache size

void matrix_mult_blocked(float *C, float *A, float *B, int N) {
    for (int i = 0; i < N; i += BLOCK)
    for (int k = 0; k < N; k += BLOCK)
    for (int j = 0; j < N; j += BLOCK)
    // Inner block fits in L1 cache
    for (int ii = i; ii < i + BLOCK && ii < N; ii++)
    for (int kk = k; kk < k + BLOCK && kk < N; kk++)
    for (int jj = j; jj < j + BLOCK && jj < N; jj++)
        C[ii*N+jj] += A[ii*N+kk] * B[kk*N+jj];
}

For perf cache event reference and false sharing detection patterns, see references/cache-counters.md.

  • Use skills/profilers/linux-perf for perf stat and perf record cache measurements
  • Use skills/profilers/valgrind — cachegrind simulates cache behaviour
  • Use skills/low-level-programming/simd-intrinsics — SoA layout pairs with SIMD vectorization
  • Use skills/low-level-programming/memory-model for false sharing in concurrent contexts

© mohitmishra786, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (references) in skills/low-level-programming/cpu-cache-opt of mohitmishra786/low-level-dev-skills.

  • SKILL.md
  • references/cache-counters.md

Open the folder on GitHubat commit bdc5847

Compare with similar skills

Cpu Cache Opt next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Cpu Cache Opt compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Cpu Cache Opt this skillmohitmishra786/low-level-dev-skills253—~1.7kAutomated safety check: PassMIT
SeekDB Code Reviewoceanbase/seekdb3.1k—~2.1kAutomated safety check: PassApache-2.0
Add Grammarafnanenayet/diffsitter2.4k—~1.9kAutomated safety check: NotesMIT
Cppcrazyguitar/cppcheatsheet290—~1.8kAutomated safety check: PassMIT
Dbgtheodo-group/debug-that158—~2.2kAutomated safety check: PassMIT
Style Checkernoumena-labs/Sipp121—~1.4kAutomated safety check: PassApache-2.0

Similar skills

  • SeekDB Code Review

    oceanbase/seekdb

    Reviews seekdb pull requests and diffs for real defects in correctness, resources, concurrency, security and tests, reporting only Blocker or Major findings.

    3.1k GitHub stars~2.1k tokensUpdated today
    DevelopmentAuto-check passed
  • Add Grammar

    afnanenayet/diffsitter

    Step-by-step guide for adding a new tree-sitter language grammar to diffsitter.

    2.4k GitHub stars~1.9k tokensUpdated yesterday
    DevelopmentAuto-check: notes
  • Cpp

    crazyguitar/cppcheatsheet

    Comprehensive C/C++ programming reference covering everything from C11-C23 and C++11-C++23, system programming, CUDA GPU computing, debugging tools, Rust interop, and advanced topics.

    290 GitHub stars~1.8k tokensUpdated today
    DevelopmentAuto-check passed
  • Dbg

    theodo-group/debug-that

    Debug applications using the dbg CLI debugger. An agent skill from theodo-group/debug-that.

    158 GitHub stars~2.2k tokensUpdated today
    DevelopmentAuto-check passed
  • Style Checker

    noumena-labs/Sipp

    Enforces this monorepo's coding style rules by inspecting git diffs, reading .agents/skills/style-checker/references/styleguidance.md, fixing style violations, and reporting the result.

    121 GitHub stars~1.4k tokensUpdated 18 days ago
    DevelopmentAuto-check passed
  • Readable Cpp

    crazyguitar/cppcheatsheet

    Readable C/C++/Rust/CUDA code rules inspired by The Art of Readable Code.

    290 GitHub stars~6.4k tokensUpdated today
    DevelopmentAuto-check passed

More from mohitmishra786/low-level-dev-skills

All 138 skills in this repo
  • ARM and AArch64 Assembly

    mohitmishra786/low-level-dev-skills

    Guides reading and writing AArch64 and ARM Thumb assembly: compiler output, inline asm, registers, the AAPCS calling convention and NEON or SVE basics.

    253 GitHub stars~1.9k tokensUpdated 3 mo ago
    Auto-check passed
  • RISC-V Assembly Guide

    mohitmishra786/low-level-dev-skills

    Reference for RISC-V assembly on RV32 and RV64: register names and calling convention, extension naming, GCC and Clang inline asm, and QEMU with GDB debugging.

    253 GitHub stars~1.8k tokensUpdated 3 mo ago
    Auto-check passed
  • x86-64 Assembly Reference

    mohitmishra786/low-level-dev-skills

    Explains x86-64 registers, the System V AMD64 calling convention, and how to read compiler-generated or inline assembly.

    253 GitHub stars~1.5k tokensUpdated 3 mo ago
    Auto-check passed
  • Bazel for C and C++

    mohitmishra786/low-level-dev-skills

    Guides your agent through Bazel for C/C++ projects: BUILD files, Bzlmod dependencies, toolchain registration, remote execution, dependency queries and sandbox debugging.

    253 GitHub stars~1.5k tokensUpdated 3 mo ago
    Auto-check passed
  • Binary Hardening

    mohitmishra786/low-level-dev-skills

    Binary hardening skill for security-hardened C/C++ builds. An agent skill from mohitmishra786/low-level-dev-skills.

    253 GitHub stars~2k tokensUpdated 3 mo ago
    Auto-check passed
  • Binutils

    mohitmishra786/low-level-dev-skills

    GNU binutils skill for binary manipulation and analysis. An agent skill from mohitmishra786/low-level-dev-skills.

    253 GitHub stars~1.2k tokensUpdated 3 mo ago
    Auto-check passed

Works with

Categories

Questions about Cpu Cache Opt

What does Cpu Cache Opt do?

CPU cache optimization skill for C/C++ and Rust. An agent skill from mohitmishra786/low-level-dev-skills. Cpu Cache Opt is an agent skill from mohitmishra786/low-level-dev-skills. CPU cache optimization skill for C/C++ and Rust.

When should I use Cpu Cache Opt?

Cpu Cache Opt fits situations like: diagnosing cache misses; improving data layout for cache efficiency; using perf stat cache counters; understanding false sharing.

How do I install Cpu Cache Opt in Claude Code?

Run `npx skills add mohitmishra786/low-level-dev-skills --skill cpu-cache-opt -a claude-code`. Or copy the skill folder (skills/low-level-programming/cpu-cache-opt in mohitmishra786/low-level-dev-skills) into .claude/skills/cpu-cache-opt in your project. Claude Code loads it when a task matches its description.

How do I install Cpu Cache Opt in Codex?

Run `npx skills add mohitmishra786/low-level-dev-skills --skill cpu-cache-opt -a codex`. Or copy the skill folder (skills/low-level-programming/cpu-cache-opt in mohitmishra786/low-level-dev-skills) into .agents/skills/cpu-cache-opt in your project. Codex loads it when a task matches its description.

Can I use Cpu Cache Opt in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mohitmishra786/low-level-dev-skills --skill cpu-cache-opt -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/cpu-cache-opt, .gemini/skills/cpu-cache-opt, .github/skills/cpu-cache-opt and .opencode/skills/cpu-cache-opt in your project.

What does Cpu Cache Opt need to run?

SKILL.md names no scripts, command-line tools or credentials: Cpu Cache Opt is instructions for the agent only.

Does Cpu Cache Opt access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Cpu Cache Opt safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Cpu Cache Opt use?

Cpu Cache Opt is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Cpu Cache Opt use?

About 1.7k tokens (SKILL.md is roughly 6.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.1k tokens, read only when the agent opens those files.

What are the alternatives to Cpu Cache Opt?

Skills that share tags, products or a category with Cpu Cache Opt: SeekDB Code Review (oceanbase/seekdb, 3.1k stars), Add Grammar (afnanenayet/diffsitter, 2.4k stars), Cpp (crazyguitar/cppcheatsheet, 290 stars) and Dbg (theodo-group/debug-that, 158 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Cpu Cache Opt?

mohitmishra786 (a GitHub user) maintains it in mohitmishra786/low-level-dev-skills, which has 253 GitHub stars. The repository holds 138 skills in this directory. The repository was last updated on June 27, 2026.

Source: mohitmishra786/low-level-dev-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.