Agent skill

Hardware Counters

by mohitmishra786 in mohitmishra786/low-level-dev-skills

Hardware performance counter skill for low-level CPU analysis.

MITAuto-check passedDevelopment

Install Hardware Counters

skills CLI
$ npx skills add mohitmishra786/low-level-dev-skills --skill hardware-counters -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install mohitmishra786/low-level-dev-skills hardware-counters --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/mohitmishra786/low-level-dev-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/profilers/hardware-counters .claude/skills/hardware-counters && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
hardware-counters
GitHub stars
253
Token cost
~1.7k tokens
SKILL.md length
276 words
Files
1
Skills in repo
138
Repo updated
First seen
Licence
MIT

At a glance

Hardware performance counter skill for low-level CPU analysis.

  • Works in 7 steps: perf stat — basic counter collection → Specifying PMU events with -e → Key metrics and thresholds → …
  • Collecting PMU events with perf stat
  • SKILL.md covers Purpose, Triggers, Workflow and Related skills
  • Calls cmake, pip and git; reaches github.com

What it does

Hardware Counters is an agent skill from mohitmishra786/low-level-dev-skills. Hardware performance counter skill for low-level CPU analysis. Use when collecting PMU events with perf stat, using the PAPI library, measuring cache miss rates and branch misprediction ratios, computing IPC, or correlating PMU events to source lines. Activates on queries about hardware counters, PMU events, perf stat -e, PAPI, cache miss rate, branch misprediction, IPC measurement, or CPU performance events.

Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Development, covering Performance optimization. The repository describes itself as: A curated suite of AI agent skills for systems and low-level programming with C/C++, Rust, and Zig toolchains, covering compilers, debuggers, profilers, build systems…. The licence is MIT.

When your agent uses it

  • Collecting PMU events with perf stat
  • Using the PAPI library
  • Measuring cache miss rates and branch misprediction ratios
  • Correlating PMU events to source lines

Example prompts

  • “/hardware-counters”

Requirements

  • Python 3

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. perf stat — basic counter collection
  2. Specifying PMU events with -e
  3. Key metrics and thresholds
  4. Raw PMU events (CPU-specific)
  5. Source-level annotation with perf record/annotate
  6. PAPI — Portable API for hardware counters
  7. Intel PCM (Performance Counter Monitor)

What it can do on your machine

Read from SKILL.md and the folder at commit bdc5847. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • cmake
    • pip
    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Hardware Counters loads about 1.7k tokens when it runs. Until then it costs about 108 tokens; SKILL.md has 276 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~108
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from mohitmishra786/low-level-dev-skills at commit bdc5847, republished under its MIT licence (© mohitmishra786). 276 words, ~1,654 tokens.

Download SKILL.mdSave it as .claude/skills/hardware-counters/SKILL.md (or your agent's skills folder).
name
hardware-counters
description
Hardware performance counter skill for low-level CPU analysis. Use when collecting PMU events with perf stat, using the PAPI library, measuring cache miss rates and branch misprediction ratios, computing IPC, or correlating PMU events to source lines. Activates on queries about hardware counters, PMU events, perf stat -e, PAPI, cache miss rate, branch misprediction, IPC measurement, or CPU performance events.

Hardware Performance Counters

Purpose

Guide agents through hardware performance counter analysis: collecting PMU events with perf stat -e, using the PAPI library for portable counter access, interpreting cache miss rates and branch misprediction ratios, computing IPC, and correlating events to source lines with perf annotate.

Triggers

  • "How do I measure cache miss rate with perf?"
  • "How do I count branch mispredictions?"
  • "How do I compute IPC (instructions per clock) with perf?"
  • "How do I use the PAPI library for hardware counters?"
  • "How do I see which source lines cause the most cache misses?"
  • "How do I measure memory bandwidth with performance counters?"

Workflow

1. perf stat — basic counter collection
bash
# Basic hardware event summary
perf stat ./prog

# Output:
#  Performance counter stats for './prog':
#
#      1,234,567,890      instructions
#        456,789,012      cycles
#         12,345,678      cache-misses         #    1.23 % of all cache refs
#         23,456,789      branch-misses        #    2.34 % of all branches
#
#       0.456789012 seconds time elapsed

# Derived metrics (computed from the output)
# IPC = instructions / cycles = 1,234,567,890 / 456,789,012 ≈ 2.70
# CPI = cycles / instructions ≈ 0.37
2. Specifying PMU events with -e
bash
# Specific hardware events
perf stat -e instructions,cycles,cache-misses,branch-misses ./prog

# L1/L2/L3 cache events
perf stat -e \
  L1-dcache-loads,L1-dcache-load-misses,\
  L2-loads,L2-load-misses,\
  LLC-loads,LLC-load-misses \
  ./prog

# Memory bandwidth (Intel)
perf stat -e \
  uncore_imc/cas_count_read/,\
  uncore_imc/cas_count_write/ \
  ./prog

# TLB misses
perf stat -e dTLB-loads,dTLB-load-misses,iTLB-loads,iTLB-load-misses ./prog

# Branch misprediction rate
perf stat -e branches,branch-misses ./prog
# Rate = branch-misses / branches × 100%

# Available events (varies by CPU)
perf list hardware          # generic hardware events
perf list cache             # cache events
perf list pmu               # raw PMU events for your CPU
3. Key metrics and thresholds
MetricFormulaHealthyConcerning
IPCinstructions / cycles> 2.0 (modern x86)< 1.0
L1 miss rateL1-misses / L1-accesses< 1%> 5%
LLC miss rateLLC-misses / LLC-accesses< 1%> 10%
Branch miss ratebranch-misses / branches< 1%> 5%
MPKImisses per 1K instructions—L3 MPKI > 10 = memory bound
bash
# Compute MPKI (Misses Per Kilo-Instructions)
perf stat -e instructions,LLC-load-misses ./prog
# MPKI = LLC-load-misses / (instructions / 1000)
4. Raw PMU events (CPU-specific)

For events not in the generic aliases, use raw event codes:

bash
# Intel: use perf list or look up in Intel SDM
# Format: rXXYY where XX=umask, YY=event code
perf stat -e r0124 ./prog    # example Intel raw event

# List Intel events with ocperf (OpenCL Perf Events)
pip install ocperf
ocperf.py list | grep "mem_load"

# Use libpfm4 for event names
pfm_ls | grep "MEM_LOAD"
perf stat -e $(pfm_ls | grep "MEM_LOAD_RETIRED.L3_MISS") ./prog

# AMD: similar approach
perf stat -e r04041 ./prog   # AMD raw event
5. Source-level annotation with perf record/annotate
bash
# Record with hardware events
perf record -e LLC-load-misses -g ./prog

# Annotate: show source lines sorted by cache miss count
perf annotate --stdio

# Interactive (requires debug symbols)
perf report
# Press 'a' on a function to annotate it

# Combined: record hotspot + annotate
perf record -e cycles:u -g ./prog
perf annotate --symbol=my_function --stdio 2>/dev/null | head -40

# Example annotate output:
# Percent | Source code
#   45.23 |     for (int i = 0; i < N; i++)
#    3.12 |         sum += data[i];   ← cache miss here (strided access)
6. PAPI — Portable API for hardware counters

PAPI provides a portable C API across different CPU architectures:

c
#include <papi.h>
#include <stdio.h>

int main(void) {
    int Events[] = {PAPI_TOT_INS, PAPI_TOT_CYC,
                    PAPI_L2_TCM,  PAPI_BR_MSP};
    long long values[4];

    if (PAPI_library_init(PAPI_VER_CURRENT) != PAPI_VER_CURRENT) {
        fprintf(stderr, "PAPI init failed\n");
        return 1;
    }

    PAPI_start_counters(Events, 4);

    // --- Code to measure ---
    do_work();
    // -----------------------

    PAPI_stop_counters(values, 4);

    printf("Instructions:      %lld\n", values[0]);
    printf("Cycles:            %lld\n", values[1]);
    printf("IPC:               %.2f\n", (double)values[0]/values[1]);
    printf("L2 cache misses:   %lld\n", values[2]);
    printf("Branch mispred:    %lld\n", values[3]);

    return 0;
}
bash
# Build with PAPI
gcc -O2 -g -o prog prog.c -lpapi

# Available PAPI events on your system
papi_avail -a | head -30
papi_native_avail | grep "L3"    # native events with "L3"

Common PAPI presets:

PresetEvent
PAPI_TOT_INSTotal instructions
PAPI_TOT_CYCTotal cycles
PAPI_L1_DCML1 data cache misses
PAPI_L2_TCML2 total cache misses
PAPI_L3_TCML3 total cache misses
PAPI_BR_MSPBranch mispredictions
PAPI_TLB_DMData TLB misses
PAPI_FP_INSFloating point instructions
PAPI_VEC_INSVector/SIMD instructions
7. Intel PCM (Performance Counter Monitor)
bash
# Intel PCM — system-wide counters, no root required on modern kernels
git clone https://github.com/intel/pcm
cd pcm && cmake -S . -B build && cmake --build build

# Measure memory bandwidth
./build/bin/pcm-memory 1    # sample every 1 second

# Core utilization + IPC
./build/bin/pcm 1

# Cache miss breakdown per socket
./build/bin/pcm 1 -csv | head -20
  • Use skills/profilers/intel-vtune-amd-uprof for guided microarchitecture analysis
  • Use skills/profilers/linux-perf for perf record/report and flamegraph generation
  • Use skills/low-level-programming/cpu-cache-opt for applying cache optimization patterns
  • Use skills/low-level-programming/simd-intrinsics for improving FLOPS/cycle metrics

© mohitmishra786, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/profilers/hardware-counters of mohitmishra786/low-level-dev-skills.

Open the folder on GitHubat commit bdc5847

Compare with similar skills

Hardware Counters next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Hardware Counters compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Hardware Counters this skillmohitmishra786/low-level-dev-skills253—~1.7kAutomated safety check: PassMIT
Code Review ChecklistshareAI-lab/learn-claude-code78k5 repos~1.1kAutomated safety check: PassMIT
LLM Torch Profiler Analysissgl-project/sglang37k2 repos~6.4kAutomated safety check: PassApache-2.0
Pycrazyguitar/pysheeet8.2k—~886Automated safety check: PassMIT
Cmux Debugging Guidemanaflow-ai/cmux28k1 repos~1.1kAutomated safety check: PassCustom licence
Electron Heap Snapshot Analysiskeybase/client9.3k—~875Automated safety check: PassBSD-3-Clause

Similar skills

  • Code Review Checklist

    shareAI-lab/learn-claude-code

    Reviews code against a five-part checklist covering security, correctness, performance, maintainability and testing, and reports findings in a fixed format.

    78k GitHub starsUsed in 5 repos~1.1k tokens
    DevelopmentAuto-check passed
  • LLM Torch Profiler Analysis

    sgl-project/sglang

    Unified LLM torch-profiler triage skill for sglang, vllm, TensorRT-LLM, and TokenSpeed.

    37k GitHub starsUsed in 2 repos~6.4k tokens
    DevelopmentAuto-check passed
  • Py

    crazyguitar/pysheeet

    Comprehensive Python programming reference covering syntax, concurrency, networking, databases, ML/LLM development, and HPC.

    8.2k GitHub stars~886 tokensUpdated yesterday
    DevelopmentAuto-check passed
  • Cmux Debugging Guide

    manaflow-ai/cmux

    Covers debug logging, the Debug menu, profiling rules and runtime pitfalls for working on the cmux macOS terminal app.

    28k GitHub starsUsed in 1 repo~1.1k tokens
    DevelopmentAuto-check passed
  • Analyzes V8, Chrome and Electron .heapsnapshot files with Node scripts to find memory leaks, detached DOM nodes and the retainer paths that keep objects alive.

    9.3k GitHub stars~875 tokensUpdated yesterday
    DevelopmentAuto-check passed
  • Runs controlled JMH experiments on the Caffeine cache to find shared contention and hot-path waste, then reviews correctness and returns a reviewable patch.

    18k GitHub stars~2.6k tokensUpdated 3 days ago
    DevelopmentAuto-check: notes

More from mohitmishra786/low-level-dev-skills

All 138 skills in this repo
  • ARM and AArch64 Assembly

    mohitmishra786/low-level-dev-skills

    Guides reading and writing AArch64 and ARM Thumb assembly: compiler output, inline asm, registers, the AAPCS calling convention and NEON or SVE basics.

    253 GitHub stars~1.9k tokensUpdated 3 mo ago
    Auto-check passed
  • RISC-V Assembly Guide

    mohitmishra786/low-level-dev-skills

    Reference for RISC-V assembly on RV32 and RV64: register names and calling convention, extension naming, GCC and Clang inline asm, and QEMU with GDB debugging.

    253 GitHub stars~1.8k tokensUpdated 3 mo ago
    Auto-check passed
  • x86-64 Assembly Reference

    mohitmishra786/low-level-dev-skills

    Explains x86-64 registers, the System V AMD64 calling convention, and how to read compiler-generated or inline assembly.

    253 GitHub stars~1.5k tokensUpdated 3 mo ago
    Auto-check passed
  • Bazel for C and C++

    mohitmishra786/low-level-dev-skills

    Guides your agent through Bazel for C/C++ projects: BUILD files, Bzlmod dependencies, toolchain registration, remote execution, dependency queries and sandbox debugging.

    253 GitHub stars~1.5k tokensUpdated 3 mo ago
    Auto-check passed
  • Binary Hardening

    mohitmishra786/low-level-dev-skills

    Binary hardening skill for security-hardened C/C++ builds. An agent skill from mohitmishra786/low-level-dev-skills.

    253 GitHub stars~2k tokensUpdated 3 mo ago
    Auto-check passed
  • Binutils

    mohitmishra786/low-level-dev-skills

    GNU binutils skill for binary manipulation and analysis. An agent skill from mohitmishra786/low-level-dev-skills.

    253 GitHub stars~1.2k tokensUpdated 3 mo ago
    Auto-check passed

Categories

Questions about Hardware Counters

What does Hardware Counters do?

Hardware performance counter skill for low-level CPU analysis. Hardware Counters is an agent skill from mohitmishra786/low-level-dev-skills. Hardware performance counter skill for low-level CPU analysis.

When should I use Hardware Counters?

Hardware Counters fits situations like: collecting PMU events with perf stat; using the PAPI library; measuring cache miss rates and branch misprediction ratios; correlating PMU events to source lines.

How do I install Hardware Counters in Claude Code?

Run `npx skills add mohitmishra786/low-level-dev-skills --skill hardware-counters -a claude-code`. Or copy the skill folder (skills/profilers/hardware-counters in mohitmishra786/low-level-dev-skills) into .claude/skills/hardware-counters in your project. Claude Code loads it when a task matches its description.

How do I install Hardware Counters in Codex?

Run `npx skills add mohitmishra786/low-level-dev-skills --skill hardware-counters -a codex`. Or copy the skill folder (skills/profilers/hardware-counters in mohitmishra786/low-level-dev-skills) into .agents/skills/hardware-counters in your project. Codex loads it when a task matches its description.

Can I use Hardware Counters in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mohitmishra786/low-level-dev-skills --skill hardware-counters -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/hardware-counters, .gemini/skills/hardware-counters, .github/skills/hardware-counters and .opencode/skills/hardware-counters in your project.

What does Hardware Counters need to run?

Going by SKILL.md and its folder, Hardware Counters needs the command-line tools its instructions call (cmake, pip and git). Our summary lists: Python 3.

Does Hardware Counters access the network?

SKILL.md names 1 domain. In commands or code: github.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Hardware Counters safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Hardware Counters use?

Hardware Counters is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Hardware Counters use?

About 1.7k tokens (SKILL.md is roughly 6.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Hardware Counters?

Skills that share tags, products or a category with Hardware Counters: Code Review Checklist (shareAI-lab/learn-claude-code, 78k stars), LLM Torch Profiler Analysis (sgl-project/sglang, 37k stars), Py (crazyguitar/pysheeet, 8.2k stars) and Cmux Debugging Guide (manaflow-ai/cmux, 28k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Hardware Counters?

mohitmishra786 (a GitHub user) maintains it in mohitmishra786/low-level-dev-skills, which has 253 GitHub stars. The repository holds 138 skills in this directory. The repository was last updated on June 27, 2026.

Source: mohitmishra786/low-level-dev-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.