HIP and ROCm skill for AMD GPU programming. An agent skill from mohitmishra786/low-level-dev-skills.

MITAuto-check: notesDevelopment

Install Hip Rocm

skills CLI
$ npx skills add mohitmishra786/low-level-dev-skills --skill hip-rocm -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install mohitmishra786/low-level-dev-skills hip-rocm --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/mohitmishra786/low-level-dev-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/gpu/hip-rocm .claude/skills/hip-rocm && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
hip-rocm
GitHub stars
252
Token cost
~1.6k tokens
SKILL.md length
361 words
Files
1
Skills in repo
138
Repo updated
First seen
Licence
MIT

At a glance

HIP and ROCm skill for AMD GPU programming. An agent skill from mohitmishra786/low-level-dev-skills.

  • Works in 8 steps: ROCm installation and verification → Minimal HIP kernel → CUDA → HIP porting with HIPIFY → …
  • Writing HIP kernels with hipcc
  • SKILL.md covers Purpose, When to Use, Workflow and Common Problems, plus 1 more section
  • Calls apt

What it does

Hip Rocm is an agent skill from mohitmishra786/low-level-dev-skills. HIP and ROCm skill for AMD GPU programming. Use when writing HIP kernels with hipcc, porting CUDA code via HIPIFY, profiling with rocprof, debugging with rocgdb, or optimizing for MI300X. Activates on queries about HIP, ROCm, hipify, hipcc, rocprof, or CUDA to AMD porting.

Its SKILL.md is about 1.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Development, covering Debugging and GPU and accelerator computing. It works with CUDA. The repository describes itself as: A curated suite of AI agent skills for systems and low-level programming with C/C++, Rust, and Zig toolchains, covering compilers, debuggers, profilers, build systems…. The licence is MIT.

When your agent uses it

  • Writing HIP kernels with hipcc
  • Porting CUDA code via HIPIFY
  • Profiling with rocprof
  • Debugging with rocgdb

Example prompts

  • “/hip-rocm”

Workflow steps

8 steps, taken from the step headings in SKILL.md.

  1. ROCm installation and verification
  2. Minimal HIP kernel
  3. CUDA → HIP porting with HIPIFY
  4. hipcc flags
  5. rocprof profiling
  6. rocgdb debugging
  7. MI300X optimizations
  8. Library ecosystem

What it can do on your machine

Read from SKILL.md and the folder at commit bdc5847. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • apt

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Hip Rocm loads about 1.6k tokens when it runs. Until then it costs about 71 tokens; SKILL.md has 361 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~71
When it runs · the whole SKILL.md, loaded when a task matches
~1.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteRuns commands with sudoSKILL.md:27
    sudo apt install rocm-dev rocm-libs hip-dev

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from mohitmishra786/low-level-dev-skills at commit bdc5847, republished under its MIT licence (© mohitmishra786). 361 words, ~1,581 tokens.

Download SKILL.mdSave it as .claude/skills/hip-rocm/SKILL.md (or your agent's skills folder).
name
hip-rocm
description
HIP and ROCm skill for AMD GPU programming. Use when writing HIP kernels with hipcc, porting CUDA code via HIPIFY, profiling with rocprof, debugging with rocgdb, or optimizing for MI300X. Activates on queries about HIP, ROCm, hipify, hipcc, rocprof, or CUDA to AMD porting.

HIP / ROCm

Purpose

Guide agents through AMD GPU programming with HIP: the HIP runtime API, hipcc compilation, porting CUDA code with HIPIFY (hipify-perl, hipify-clang), ROCm toolchain setup, profiling with rocprof, debugging with rocgdb, HIP-vs-CUDA API mapping, and MI300X-specific optimizations.

When to Use

  • Porting an existing CUDA codebase to AMD GPUs
  • Setting up ROCm on Linux for MI200/MI300 hardware
  • Writing native HIP kernels for AMD data center GPUs
  • Profiling HIP applications with rocprof or rocprofiler-sdk
  • Debugging device faults with rocgdb or compute sanitizers
  • Building multi-vendor GPU code with HIP portability macros

Workflow

1. ROCm installation and verification
bash
# Ubuntu/Debian (check ROCm docs for your distro version)
sudo apt install rocm-dev rocm-libs hip-dev

# Verify
rocminfo | head -30
hipconfig --version
hipcc --version

# List devices
rocm-smi

Set GPU target for compilation:

bash
export AMDGPU_TARGETS=gfx942   # MI300X
export HIP_PLATFORM=amd
2. Minimal HIP kernel
cpp
// vector_add.hip
#include <hip/hip_runtime.h>
#include <stdio.h>

__global__ void vector_add(const float *a, const float *b, float *c, int n) {
    int i = blockIdx.x * blockDim.x + threadIdx.x;
    if (i < n)
        c[i] = a[i] + b[i];
}

int main(void) {
    const int n = 1 << 20;
    size_t bytes = n * sizeof(float);
    float *d_a, *d_b, *d_c;

    hipMalloc(&d_a, bytes);
    hipMalloc(&d_b, bytes);
    hipMalloc(&d_c, bytes);

    int threads = 256;
    int blocks = (n + threads - 1) / threads;
    hipLaunchKernelGGL(vector_add, dim3(blocks), dim3(threads), 0, 0,
                       d_a, d_b, d_c, n);
    hipDeviceSynchronize();

    hipFree(d_a); hipFree(d_b); hipFree(d_c);
    return 0;
}
bash
hipcc -O3 --offload-arch=gfx942 -o vector_add vector_add.hip
./vector_add
3. CUDA → HIP porting with HIPIFY
bash
# Perl-based batch converter (quick port)
hipify-perl cuda_kernel.cu > cuda_kernel.hip

# Clang-based (more accurate, preserves structure)
hipify-clang cuda_project/ -o hip_project/ --cuda-path=/usr/local/cuda

# Convert single file in place
hipify-clang -inplace --cuda-path=/usr/local/cuda main.cu

Common API mappings:

CUDAHIP
cudaMallochipMalloc
cudaMemcpyhipMemcpy
cudaMemcpyAsynchipMemcpyAsync
cudaStream_thipStream_t
<<<grid, block>>>hipLaunchKernelGGL or <<<>>> (HIP supports CUDA syntax)
__syncthreads()__syncthreads() (same)
threadIdx / blockIdxSame builtins

Portability header for dual compilation:

cpp
#ifdef __HIP_PLATFORM_AMD__
#include <hip/hip_runtime.h>
#else
#include <cuda_runtime.h>
#define hipMalloc cudaMalloc
#define hipMemcpy cudaMemcpy
// ... more macros
#endif
4. hipcc flags
bash
# Target specific GPU architecture
hipcc --offload-arch=gfx942 -O3 -o app main.hip

# Multiple architectures
hipcc --offload-arch=gfx90a --offload-arch=gfx942 -o app main.hip

# Debug
hipcc -g -O0 --offload-arch=gfx942 -o app_debug main.hip

# Link with rocBLAS
hipcc -lrocblas -o app main.hip
5. rocprof profiling
bash
# Basic kernel trace
rocprof --stats ./app

# CSV metrics output
rocprof -i input.csv -o output.csv ./app

# input.csv example:
# pmc: SQ_INSTS_VALU_ADD_F32,SQ_INSTS_VALU_MUL_F32,GRBM_COUNT
bash
# ROCm 6.x rocprofiler-sdk (preferred for new projects)
rocprofv3 --kernel-trace -- ./app

Key metrics (AMD terminology):

  • VALU utilization — compute unit activity
  • LDS bank conflicts — shared memory (LDS) stalls
  • Memory throughput — HBM bandwidth utilization
6. rocgdb debugging
bash
# Build with debug symbols
hipcc -g -O0 --offload-arch=gfx942 -o app_debug main.hip

rocgdb ./app_debug
gdb
(rocgdb) break vector_add
(rocgdb) run
(rocgdb) info rocm kernels
(rocgdb) rocm thread 0 0 0
(rocgdb) print i

AMD also supports compute-sanitizer equivalents via ROCm's roc-obj-extract and memory checking tools where available.

7. MI300X optimizations
bash
# Enable MFMA (matrix fused multiply-add) instructions
hipcc --offload-arch=gfx942 -munsafe-fp-atomics -O3 -o app main.hip
OptimizationMI300X note
Matrix opsUse rocBLAS/hipBLASLt for GEMM; MFMA intrinsics for custom
HBM bandwidth~5.3 TB/s peak (MI300X) — maximize memory coalescing to approach it
Wavefront size64 threads (vs CUDA warp 32) — adjust reduction patterns
LDS (shared mem)64 KB per CU; watch bank conflicts

Wavefront-aware reduction:

cpp
__device__ float warp_reduce_sum(float val) {
    // AMD wavefront = 64 lanes
    for (int offset = 32; offset > 0; offset >>= 1)
        val += __shfl_down(val, offset);
    return val;
}
Show full SKILL.md (122 more words)Show less
8. Library ecosystem
NVIDIAAMD ROCm
cuBLASrocBLAS / hipBLAS
cuDNNMIOpen
NCCLrccl
ThrusthipCUB (portable)
cuFFTrocFFT
bash
hipcc -lrocblas -o gemm_test gemm.hip

Common Problems

SymptomCauseFix
hipErrorNoDeviceROCm driver not loadedCheck rocm-smi; add user to render group
Wrong architecture binaryMismatched gfx* targetrocminfo → set --offload-arch
hipify incomplete portCUDA-specific APIsManual fix: cooperative groups, texture refs
Slower than CUDA referenceWavefront 64 vs warp 32Tune block size to multiples of 64
HSA_STATUS_ERRORGPU busy or OOMrocm-smi --showmeminfo; reduce allocation
rocprof empty outputNo kernels launchedVerify hipGetLastError() after launch
  • skills/gpu/cuda — source CUDA patterns being ported
  • skills/gpu/cuda-profiling — Nsight concepts map to rocprof
  • skills/gpu/gpu-memory-model — wavefront vs warp, coalescing rules
  • skills/gpu/triton-lang — Triton supports AMD via ROCm backend
  • skills/compilers/llvm — HIP uses Clang/LLVM toolchain

© mohitmishra786, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/gpu/hip-rocm of mohitmishra786/low-level-dev-skills.

Open the folder on GitHubat commit bdc5847

Compare with similar skills

Hip Rocm next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Hip Rocm compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Hip Rocm this skillmohitmishra786/low-level-dev-skills252—~1.6kAutomated safety check: NotesMIT
Debug Distributed Hangsgl-project/sglang37k2 repos~2.4kAutomated safety check: PassApache-2.0
CUTLASS FMHA Incremental Rebuildmicrosoft/onnxruntime22k—~1.3kAutomated safety check: PassMIT
Cudatechnillogue/ptx-isa-markdown229—~2.5kAutomated safety check: PassNone
Cuda Cpp Kernelvipshop/cache-dit1.3k—~2.3kAutomated safety check: PassApache-2.0
Aoti Debugpytorch/pytorch104k1 repos~1.7kAutomated safety check: PassCustom licence

Similar skills

  • Debug Distributed Hang

    sgl-project/sglang

    Debug hanging issues in SGLang distributed inference (TP/PP/DP/EP).

    37k GitHub starsUsed in 2 repos~2.4k tokens
    DevelopmentAuto-check passed
  • Official

    Explains why editing CUTLASS fused-MHA headers in ONNX Runtime can leave stale CUDA kernels after an incremental build, and how to force and verify a real rebuild.

    22k GitHub stars~1.3k tokensUpdated today
    DevelopmentAuto-check passed
  • Cuda

    technillogue/ptx-isa-markdown

    CUDA kernel development, debugging, and performance optimization for Claude Code.

    229 GitHub stars~2.5k tokensUpdated 9 mo ago
    DevelopmentAuto-check passed
  • Cuda Cpp Kernel

    vipshop/cache-dit

    A skill your agent uses when writing, debugging, porting, reviewing, or optimizing CUDA C++ or PTX kernels; investigating CUDA Runtime or Driver API behavior; profiling kernels with Nsight Systems…

    1.3k GitHub stars~2.3k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Aoti Debug

    pytorch/pytorch

    Debug AOTInductor (AOTI) errors and crashes. An agent skill from pytorch/pytorch.

    104k GitHub starsUsed in 1 repo~1.7k tokens
    DevelopmentAuto-check passed
  • The Art of Debugging

    stas00/the-art-of-debugging

    Condensed debugging method and tool recipes for Unix, Python and PyTorch programs: crashes, hangs, segfaults, wrong output, CUDA OOM, NaN values and slowness.

    1.7k GitHub stars~6.1k tokensUpdated 4 days ago
    DevelopmentAuto-check: notes

More from mohitmishra786/low-level-dev-skills

All 138 skills in this repo
  • ARM and AArch64 Assembly

    mohitmishra786/low-level-dev-skills

    Guides reading and writing AArch64 and ARM Thumb assembly: compiler output, inline asm, registers, the AAPCS calling convention and NEON or SVE basics.

    252 GitHub stars~1.9k tokensUpdated 3 mo ago
    Auto-check passed
  • RISC-V Assembly Guide

    mohitmishra786/low-level-dev-skills

    Reference for RISC-V assembly on RV32 and RV64: register names and calling convention, extension naming, GCC and Clang inline asm, and QEMU with GDB debugging.

    252 GitHub stars~1.8k tokensUpdated 3 mo ago
    Auto-check passed
  • x86-64 Assembly Reference

    mohitmishra786/low-level-dev-skills

    Explains x86-64 registers, the System V AMD64 calling convention, and how to read compiler-generated or inline assembly.

    252 GitHub stars~1.5k tokensUpdated 3 mo ago
    Auto-check passed
  • Bazel for C and C++

    mohitmishra786/low-level-dev-skills

    Guides your agent through Bazel for C/C++ projects: BUILD files, Bzlmod dependencies, toolchain registration, remote execution, dependency queries and sandbox debugging.

    252 GitHub stars~1.5k tokensUpdated 3 mo ago
    Auto-check passed
  • Binary Hardening

    mohitmishra786/low-level-dev-skills

    Binary hardening skill for security-hardened C/C++ builds. An agent skill from mohitmishra786/low-level-dev-skills.

    252 GitHub stars~2k tokensUpdated 3 mo ago
    Auto-check passed
  • Binutils

    mohitmishra786/low-level-dev-skills

    GNU binutils skill for binary manipulation and analysis. An agent skill from mohitmishra786/low-level-dev-skills.

    252 GitHub stars~1.2k tokensUpdated 3 mo ago
    Auto-check passed

Works with

Categories

Questions about Hip Rocm

What does Hip Rocm do?

HIP and ROCm skill for AMD GPU programming. An agent skill from mohitmishra786/low-level-dev-skills. Hip Rocm is an agent skill from mohitmishra786/low-level-dev-skills. HIP and ROCm skill for AMD GPU programming.

When should I use Hip Rocm?

Hip Rocm fits situations like: writing HIP kernels with hipcc; porting CUDA code via HIPIFY; profiling with rocprof; debugging with rocgdb.

How do I install Hip Rocm in Claude Code?

Run `npx skills add mohitmishra786/low-level-dev-skills --skill hip-rocm -a claude-code`. Or copy the skill folder (skills/gpu/hip-rocm in mohitmishra786/low-level-dev-skills) into .claude/skills/hip-rocm in your project. Claude Code loads it when a task matches its description.

How do I install Hip Rocm in Codex?

Run `npx skills add mohitmishra786/low-level-dev-skills --skill hip-rocm -a codex`. Or copy the skill folder (skills/gpu/hip-rocm in mohitmishra786/low-level-dev-skills) into .agents/skills/hip-rocm in your project. Codex loads it when a task matches its description.

Can I use Hip Rocm in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mohitmishra786/low-level-dev-skills --skill hip-rocm -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/hip-rocm, .gemini/skills/hip-rocm, .github/skills/hip-rocm and .opencode/skills/hip-rocm in your project.

What does Hip Rocm need to run?

Going by SKILL.md and its folder, Hip Rocm needs the command-line tools its instructions call (apt).

Does Hip Rocm access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Hip Rocm safe to install?

Our automated static check of SKILL.md found notes only (runs commands with sudo), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Hip Rocm use?

Hip Rocm is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Hip Rocm use?

About 1.6k tokens (SKILL.md is roughly 6.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Hip Rocm?

Skills that share tags, products or a category with Hip Rocm: Debug Distributed Hang (sgl-project/sglang, 37k stars), CUTLASS FMHA Incremental Rebuild (microsoft/onnxruntime, 22k stars), Cuda (technillogue/ptx-isa-markdown, 229 stars) and Cuda Cpp Kernel (vipshop/cache-dit, 1.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Hip Rocm?

mohitmishra786 (a GitHub user) maintains it in mohitmishra786/low-level-dev-skills, which has 252 GitHub stars. The repository holds 138 skills in this directory. The repository was last updated on June 27, 2026.

Source: mohitmishra786/low-level-dev-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.