Agent skill

Intel Vtune Amd Uprof

by mohitmishra786 in mohitmishra786/low-level-dev-skills

Intel VTune and AMD uProf profiling skill for microarchitecture analysis.

MITAuto-check passedDevelopment

Install Intel Vtune Amd Uprof

skills CLI
$ npx skills add mohitmishra786/low-level-dev-skills --skill intel-vtune-amd-uprof -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install mohitmishra786/low-level-dev-skills intel-vtune-amd-uprof --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/mohitmishra786/low-level-dev-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/profilers/intel-vtune-amd-uprof .claude/skills/intel-vtune-amd-uprof && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
intel-vtune-amd-uprof
GitHub stars
252
Token cost
~1.7k tokens
SKILL.md length
357 words
Files
1
Skills in repo
138
Repo updated
First seen
Licence
MIT

At a glance

Intel VTune and AMD uProf profiling skill for microarchitecture analysis.

  • Works in 7 steps: VTune setup (free Community Edition) → Analysis types → Hotspot analysis → …
  • Analyzing hotspots
  • SKILL.md covers Purpose, Triggers, Workflow and Related skills
  • Reaches intel.com and amd.com

What it does

Intel Vtune Amd Uprof is an agent skill from mohitmishra786/low-level-dev-skills. Intel VTune and AMD uProf profiling skill for microarchitecture analysis. Use when analyzing hotspots, microarchitecture bottlenecks, memory access patterns, pipeline stalls, or using the roofline model. Covers VTune Community Edition (free) and AMD uProf as a free alternative. Activates on queries about VTune, uProf, microarchitecture analysis, pipeline stalls, memory bandwidth, roofline model, or hardware performance analysis.

Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Development, covering Performance optimization. The repository describes itself as: A curated suite of AI agent skills for systems and low-level programming with C/C++, Rust, and Zig toolchains, covering compilers, debuggers, profilers, build systems…. The licence is MIT.

When your agent uses it

  • Analyzing hotspots
  • Microarchitecture bottlenecks
  • Memory access patterns
  • Pipeline stalls

Example prompts

  • “/intel-vtune-amd-uprof”

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. VTune setup (free Community Edition)
  2. Analysis types
  3. Hotspot analysis
  4. Microarchitecture exploration — pipeline stalls
  5. Memory access analysis
  6. AMD uProf — free alternative for AMD CPUs
  7. Roofline model

What it can do on your machine

Read from SKILL.md and the folder at commit bdc5847. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are bash).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • intel.com
    • amd.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Intel Vtune Amd Uprof loads about 1.7k tokens when it runs. Until then it costs about 114 tokens; SKILL.md has 357 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~114
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from mohitmishra786/low-level-dev-skills at commit bdc5847, republished under its MIT licence (© mohitmishra786). 357 words, ~1,673 tokens.

Download SKILL.mdSave it as .claude/skills/intel-vtune-amd-uprof/SKILL.md (or your agent's skills folder).
name
intel-vtune-amd-uprof
description
Intel VTune and AMD uProf profiling skill for microarchitecture analysis. Use when analyzing hotspots, microarchitecture bottlenecks, memory access patterns, pipeline stalls, or using the roofline model. Covers VTune Community Edition (free) and AMD uProf as a free alternative. Activates on queries about VTune, uProf, microarchitecture analysis, pipeline stalls, memory bandwidth, roofline model, or hardware performance analysis.

Intel VTune & AMD uProf

Purpose

Guide agents through CPU microarchitecture profiling with Intel VTune Profiler (free Community Edition) and AMD uProf: hotspot identification, microarchitecture analysis, memory access pattern optimization, pipeline stall diagnosis, and roofline model analysis.

Triggers

  • "How do I use Intel VTune to profile my code?"
  • "What are pipeline stalls and how do I reduce them?"
  • "How do I analyze memory bandwidth with VTune?"
  • "What is the roofline model and how do I use it?"
  • "How do I use AMD uProf as a free alternative to VTune?"
  • "My code has good cache hit rates but is still slow"

Workflow

1. VTune setup (free Community Edition)
bash
# Download Intel VTune Profiler (Community Edition — free)
# https://www.intel.com/content/www/us/en/developer/tools/oneapi/vtune-profiler.html

# Install on Linux
source /opt/intel/oneapi/vtune/latest/env/vars.sh

# CLI usage
vtune -collect hotspots ./prog
vtune -collect microarchitecture-exploration ./prog
vtune -collect memory-access ./prog

# View results in GUI
vtune-gui &
# File → Open Result → select .vtune directory

# Or use amplxe-cl (legacy CLI)
amplxe-cl -collect hotspots ./prog
amplxe-cl -report hotspots -r result/
2. Analysis types
AnalysisWhat it findsWhen to use
HotspotsCPU-bound functionsFirst step — find where time is spent
Microarchitecture ExplorationIPC, pipeline stalls, retired instructionsAfter hotspot — why is the hotspot slow?
Memory AccessCache misses, DRAM bandwidth, NUMAMemory-bound code
ThreadingLock contention, parallel efficiencyMultithreaded code
HPC PerformanceVectorization, memory, rooflineHPC / scientific code
I/ODisk and network bottlenecksI/O-bound code
3. Hotspot analysis
bash
# Collect and report hotspots
vtune -collect hotspots -result-dir hotspots_result ./prog

# Report top functions by CPU time
vtune -report hotspots -r hotspots_result -format csv | head -20

# CLI output example:
# Function       CPU Time  Module
# compute_fft    4.532s    libfft.so
# matrix_mult    2.108s    prog
# parse_input    0.234s    prog

Build with debug info for meaningful symbols:

bash
gcc -O2 -g ./prog.c -o prog     # symbols visible in VTune
gcc -O2 -g -gsplit-dwarf -fno-omit-frame-pointer ./prog.c -o prog  # better stacks
Show full SKILL.md (176 more words)Show less
4. Microarchitecture exploration — pipeline stalls
bash
vtune -collect microarchitecture-exploration -r micro_result ./prog
vtune -report summary -r micro_result

Key metrics to examine:

MetricMeaningGood value
IPC (Instructions Per Clock)How many instructions retire per cyclex86: aim for > 2.0
CPI (Clocks Per Instruction)Inverse of IPCLower is better
Bad SpeculationBranch mispredictions< 5%
Front-End BoundInstruction decode bottleneck< 15%
Back-End BoundExecution unit or memory stall< 30%
RetiringUseful work fraction> 70% ideal
Memory Bound% cycles waiting for memory< 20%
Pipeline Analysis (Top-Down Methodology):
├── Retiring (good, useful work)
├── Bad Speculation (branch mispredictions)
├── Front-End Bound
│   ├── Fetch Latency (I-cache misses, branch mispredicts)
│   └── Fetch Bandwidth
└── Back-End Bound
    ├── Memory Bound
    │   ├── L1 Bound → L1 cache misses
    │   ├── L2 Bound → L2 cache misses
    │   ├── L3 Bound → L3 cache misses
    │   └── DRAM Bound → main memory bandwidth limited
    └── Core Bound → ALU/compute bound
5. Memory access analysis
bash
# Collect memory access profile
vtune -collect memory-access -r mem_result ./prog

# Key output sections:
# - Memory Bound: % time waiting for memory
# - LLC (Last Level Cache) Miss Rate
# - DRAM Bandwidth: GB/s achieved vs theoretical peak
# - NUMA: cross-socket accesses (for multi-socket systems)

Reading DRAM bandwidth:

DRAM Bandwidth: 18.4 GB/s
Peak Theoretical: 51.2 GB/s
Utilization: 36% — likely not DRAM-bound

If DRAM-bound: optimize data layout (AoS → SoA), reduce working set, improve spatial locality.

6. AMD uProf — free alternative for AMD CPUs
bash
# Download AMD uProf
# https://www.amd.com/en/developer/uprof.html

# CLI profiling
AMDuProfCLI collect --config tbp ./prog          # time-based profiling
AMDuProfCLI collect --config assess ./prog       # microarchitecture assessment
AMDuProfCLI collect --config memory ./prog       # memory access

# Generate report
AMDuProfCLI report -i /tmp/uprof_result/ -o report.html

# Open GUI
AMDuProf &

AMD uProf metrics map to VTune equivalents:

  • Retired Instructions → IPC analysis
  • Branch Mispredictions → Bad Speculation
  • L1/L2/L3 Cache Misses → Memory Bound levels
  • Data Cache Accesses → Cache efficiency
7. Roofline model

The roofline model shows whether code is compute-bound or memory-bound by comparing achieved performance against hardware limits:

Performance (GFLOPS/s)
     |                    _______________
Peak |                 /
Perf |              /  compute bound
     |           /
     |        /
     |     /  memory bandwidth bound
     |  /
     +------------------------------→
        Arithmetic Intensity (FLOPS/Byte)
bash
# VTune roofline collection
vtune -collect hpc-performance -r roofline_result ./prog
# Then: VTune GUI → Roofline view

# For manual calculation:
# Arithmetic Intensity = FLOPS / memory_bytes_accessed
# Peak FLOPS = CPUs × cores × freq × FLOPS_per_cycle_per_core
# Peak BW = from hardware spec (e.g., 51.2 GB/s for DDR4-3200 dual channel)

# likwid-perfctr for manual roofline data (Linux)
likwid-perfctr -C 0 -g FLOPS_DP ./prog          # double-precision FLOPS
likwid-perfctr -C 0 -g MEM ./prog               # memory bandwidth
  • Use skills/profilers/hardware-counters for raw PMU event collection with perf stat
  • Use skills/profilers/linux-perf for perf-based profiling on Linux
  • Use skills/low-level-programming/cpu-cache-opt for memory access pattern optimization
  • Use skills/low-level-programming/simd-intrinsics for vectorization to increase FLOPS

© mohitmishra786, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/profilers/intel-vtune-amd-uprof of mohitmishra786/low-level-dev-skills.

Open the folder on GitHubat commit bdc5847

Compare with similar skills

Intel Vtune Amd Uprof next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Intel Vtune Amd Uprof compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Intel Vtune Amd Uprof this skillmohitmishra786/low-level-dev-skills252—~1.7kAutomated safety check: PassMIT
Code Review ChecklistshareAI-lab/learn-claude-code78k4 repos~1.1kAutomated safety check: PassMIT
LLM Torch Profiler Analysissgl-project/sglang37k2 repos~6.4kAutomated safety check: PassApache-2.0
Pycrazyguitar/pysheeet8.2k—~886Automated safety check: PassMIT
Cmux Debugging Guidemanaflow-ai/cmux28k1 repos~1.1kAutomated safety check: PassCustom licence
Analyzing .NET Performancedotnet/skills5.6k3 repos~3.1kAutomated safety check: PassMIT

Similar skills

  • Code Review Checklist

    shareAI-lab/learn-claude-code

    Reviews code against a five-part checklist covering security, correctness, performance, maintainability and testing, and reports findings in a fixed format.

    78k GitHub starsUsed in 4 repos~1.1k tokens
    DevelopmentAuto-check passed
  • LLM Torch Profiler Analysis

    sgl-project/sglang

    Unified LLM torch-profiler triage skill for sglang, vllm, TensorRT-LLM, and TokenSpeed.

    37k GitHub starsUsed in 2 repos~6.4k tokens
    DevelopmentAuto-check passed
  • Py

    crazyguitar/pysheeet

    Comprehensive Python programming reference covering syntax, concurrency, networking, databases, ML/LLM development, and HPC.

    8.2k GitHub stars~886 tokensUpdated 3 days ago
    DevelopmentAuto-check passed
  • Cmux Debugging Guide

    manaflow-ai/cmux

    Covers debug logging, the Debug menu, profiling rules and runtime pitfalls for working on the cmux macOS terminal app.

    28k GitHub starsUsed in 1 repo~1.1k tokens
    DevelopmentAuto-check passed
  • Official

    Scans C# and .NET code for about 50 performance anti-patterns and reports prioritized findings with concrete fixes, at a scan depth you choose.

    5.6k GitHub starsUsed in 3 repos~3.1k tokens
    DevelopmentAuto-check passed
  • Analyzes V8, Chrome and Electron .heapsnapshot files with Node scripts to find memory leaks, detached DOM nodes and the retainer paths that keep objects alive.

    9.3k GitHub stars~875 tokensUpdated yesterday
    DevelopmentAuto-check passed

More from mohitmishra786/low-level-dev-skills

All 138 skills in this repo
  • ARM and AArch64 Assembly

    mohitmishra786/low-level-dev-skills

    Guides reading and writing AArch64 and ARM Thumb assembly: compiler output, inline asm, registers, the AAPCS calling convention and NEON or SVE basics.

    252 GitHub stars~1.9k tokensUpdated 3 mo ago
    Auto-check passed
  • RISC-V Assembly Guide

    mohitmishra786/low-level-dev-skills

    Reference for RISC-V assembly on RV32 and RV64: register names and calling convention, extension naming, GCC and Clang inline asm, and QEMU with GDB debugging.

    252 GitHub stars~1.8k tokensUpdated 3 mo ago
    Auto-check passed
  • x86-64 Assembly Reference

    mohitmishra786/low-level-dev-skills

    Explains x86-64 registers, the System V AMD64 calling convention, and how to read compiler-generated or inline assembly.

    252 GitHub stars~1.5k tokensUpdated 3 mo ago
    Auto-check passed
  • Bazel for C and C++

    mohitmishra786/low-level-dev-skills

    Guides your agent through Bazel for C/C++ projects: BUILD files, Bzlmod dependencies, toolchain registration, remote execution, dependency queries and sandbox debugging.

    252 GitHub stars~1.5k tokensUpdated 3 mo ago
    Auto-check passed
  • Binary Hardening

    mohitmishra786/low-level-dev-skills

    Binary hardening skill for security-hardened C/C++ builds. An agent skill from mohitmishra786/low-level-dev-skills.

    252 GitHub stars~2k tokensUpdated 3 mo ago
    Auto-check passed
  • Binutils

    mohitmishra786/low-level-dev-skills

    GNU binutils skill for binary manipulation and analysis. An agent skill from mohitmishra786/low-level-dev-skills.

    252 GitHub stars~1.2k tokensUpdated 3 mo ago
    Auto-check passed

Categories

Questions about Intel Vtune Amd Uprof

What does Intel Vtune Amd Uprof do?

Intel VTune and AMD uProf profiling skill for microarchitecture analysis. Intel Vtune Amd Uprof is an agent skill from mohitmishra786/low-level-dev-skills. Intel VTune and AMD uProf profiling skill for microarchitecture analysis.

When should I use Intel Vtune Amd Uprof?

Intel Vtune Amd Uprof fits situations like: analyzing hotspots; microarchitecture bottlenecks; memory access patterns; pipeline stalls.

How do I install Intel Vtune Amd Uprof in Claude Code?

Run `npx skills add mohitmishra786/low-level-dev-skills --skill intel-vtune-amd-uprof -a claude-code`. Or copy the skill folder (skills/profilers/intel-vtune-amd-uprof in mohitmishra786/low-level-dev-skills) into .claude/skills/intel-vtune-amd-uprof in your project. Claude Code loads it when a task matches its description.

How do I install Intel Vtune Amd Uprof in Codex?

Run `npx skills add mohitmishra786/low-level-dev-skills --skill intel-vtune-amd-uprof -a codex`. Or copy the skill folder (skills/profilers/intel-vtune-amd-uprof in mohitmishra786/low-level-dev-skills) into .agents/skills/intel-vtune-amd-uprof in your project. Codex loads it when a task matches its description.

Can I use Intel Vtune Amd Uprof in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mohitmishra786/low-level-dev-skills --skill intel-vtune-amd-uprof -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/intel-vtune-amd-uprof, .gemini/skills/intel-vtune-amd-uprof, .github/skills/intel-vtune-amd-uprof and .opencode/skills/intel-vtune-amd-uprof in your project.

What does Intel Vtune Amd Uprof need to run?

SKILL.md names no scripts, command-line tools or credentials: Intel Vtune Amd Uprof is instructions for the agent only.

Does Intel Vtune Amd Uprof access the network?

SKILL.md names 2 domains. In commands or code: intel.com and amd.com; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.

Is Intel Vtune Amd Uprof safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Intel Vtune Amd Uprof use?

Intel Vtune Amd Uprof is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Intel Vtune Amd Uprof use?

About 1.7k tokens (SKILL.md is roughly 6.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Intel Vtune Amd Uprof?

Skills that share tags, products or a category with Intel Vtune Amd Uprof: Code Review Checklist (shareAI-lab/learn-claude-code, 78k stars), LLM Torch Profiler Analysis (sgl-project/sglang, 37k stars), Py (crazyguitar/pysheeet, 8.2k stars) and Cmux Debugging Guide (manaflow-ai/cmux, 28k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Intel Vtune Amd Uprof?

mohitmishra786 (a GitHub user) maintains it in mohitmishra786/low-level-dev-skills, which has 252 GitHub stars. The repository holds 138 skills in this directory. The repository was last updated on June 27, 2026.

Source: mohitmishra786/low-level-dev-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.