Agent skill

Magpie Kernel Evaluator

by amd in amd/skills

Benchmarks LLM inference and drives GPU kernel optimization with Magpie.

MITAuto-check passedAI & LLM Engineering

Install Magpie Kernel Evaluator

skills CLI
$ npx skills add amd/skills --skill magpie-kernel-evaluator -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install amd/skills magpie-kernel-evaluator --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/amd/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/magpie-kernel-evaluator .claude/skills/magpie-kernel-evaluator && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
magpie-kernel-evaluator
GitHub stars
398
Token cost
~2.3k tokens
SKILL.md length
1,046 words
Files
10
Skills in repo
9
Repo updated
First seen
Licence
MIT

At a glance

Benchmarks LLM inference and drives GPU kernel optimization with Magpie.

  • Works in 3 steps: Benchmark an inference workload and… → Analyze or compare GPU kernels for… → Drive an optimization loop from a…
  • The user wants to benchmark vLLM
  • SKILL.md covers Choose the workflow, Preflight, Analyze a kernel and Compare kernel variants, plus 6 more sections
  • Calls python; reaches github.com

What it does

Magpie Kernel Evaluator is an agent skill from amd/skills. Benchmarks LLM inference and drives GPU kernel optimization with Magpie. Use when the user wants to benchmark vLLM, SGLang, or Atom; capture torch traces; post-process inference traces with TraceLens into prefill/decode and roofline reports; identify top bottleneck kernels or map profiler names to source; analyze or compare HIP, CUDA, PyTorch, or Triton kernels; validate and rank optimized variants; run local, container, or Ray workloads; or mentions Magpie, TraceLens, gap analysis, TTFT, TPOT, kernel evaluation…

Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 17 other files (for example `.federated.json`, `evals/evals.json` and `evals/files/analyze-simple-hip/examples/simple_hip_test/analyze_default.yaml`).

It sits in AI & LLM Engineering, covering GPU and accelerator computing, LLM inference and serving and Performance optimization. It works with SGLang, vLLM, CUDA and PyTorch. The repository describes itself as: Official AMD catalog of AI agent skills. Empower your AI agents with AMD's optimized SW stack. The licence is MIT.

When your agent uses it

  • The user wants to benchmark vLLM
  • Capture torch traces
  • Post-process inference traces with TraceLens into prefill/decode and roofline reports
  • Identify top bottleneck kernels

Example prompts

  • “Use the magpie-kernel-evaluator skill to benchmark LLM inference and drives GPU kernel optimization with Magpie”
  • “/magpie-kernel-evaluator”

Requirements

  • Python 3
  • Docker

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Benchmark an inference workload and collect throughput, latency, and traces.
  2. Analyze or compare GPU kernels for correctness and performance.
  3. Drive an optimization loop from a benchmark bottleneck to source, candidate kernels, and end-to-end validation.

What it can do on your machine

Read from SKILL.md and the folder at commit 6c92b41. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Magpie Kernel Evaluator loads about 2.3k tokens when it runs. Until then it costs about 142 tokens; SKILL.md has 1,046 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~142
When it runs · the whole SKILL.md, loaded when a task matches
~2.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from amd/skills at commit 6c92b41, republished under its MIT licence (© amd). 1,046 words, ~2,333 tokens.

Download SKILL.mdSave it as .claude/skills/magpie-kernel-evaluator/SKILL.md (or your agent's skills folder). This skill also uses 9 other files; get the full folder from GitHub.
name
magpie-kernel-evaluator
description
Benchmarks LLM inference and drives GPU kernel optimization with Magpie. Use when the user wants to benchmark vLLM, SGLang, or Atom; capture torch traces; post-process inference traces with TraceLens into prefill/decode and roofline reports; identify top bottleneck kernels or map profiler names to source; analyze or compare HIP, CUDA, PyTorch, or Triton kernels; validate and rank optimized variants; run local, container, or Ray workloads; or mentions Magpie, TraceLens, gap analysis, TTFT, TPOT, kernel evaluation, or AMD GPU optimization.

Magpie

Use Magpie for three connected jobs:

  1. Benchmark an inference workload and collect throughput, latency, and traces.
  2. Analyze or compare GPU kernels for correctness and performance.
  3. Drive an optimization loop from a benchmark bottleneck to source, candidate kernels, and end-to-end validation.

Describe only capabilities supported by the checked-out Magpie version. Do not infer support for an unverified ROCm, GPU, framework, or experimental integration.

Choose the workflow

User goalWorkflow
Evaluate one implementationanalyze
Rank two or more implementationscompare
Measure model-serving performancebenchmark
Find expensive kernels in existing tracesstandalone gap analysis
Explain a profiled inference workloadbenchmark → TraceLens post-processing → stage/roofline review
Optimize an end-to-end workloadbenchmark → TraceLens/gap analysis → source mapping → analyze/compare → re-benchmark

Use a YAML config for reproducible or multi-step work. Use inline CLI arguments for small exploratory runs.

Preflight

Installing this skill does not install the Magpie application (magpie-eval). For a preparation-only request, read the supplied fixtures and reference.md, write the requested plan or configuration, and identify commands that remain unverified. Do not install packages or execute a workload when the user forbids it; a missing Magpie installation does not prevent preparing a plan.

Before executing a Magpie workload:

  1. Check magpie --help in the intended Python environment (Python 3.10+). If the command is missing, try python -m Magpie --help in that same environment; a working module entry point can be used instead of the CLI. No module named Magpie means the application is not installed in that environment. For errors inside an installed Magpie package, diagnose the reported import or dependency failure instead.

  2. If Magpie is not installed, install it in the intended environment before continuing:

    bash
    python -m pip install git+https://github.com/AMD-AGI/Magpie.git

    For an existing Magpie source checkout, use python -m pip install -e /path/to/Magpie instead. The installed skill folder and a testcase workspace are not Magpie source checkouts. Re-run magpie --help or python -m Magpie --help to verify the installation; if it fails, report the error before attempting a workload.

  3. Inspect the installed interface for the selected workflow (substitute python -m Magpie if using the module entry point):

    bash
    magpie --help
    magpie analyze --help
    magpie compare --help
    magpie benchmark --help
    magpie --gpu-info
  4. Check required tools, model access, GPU visibility, writable output space, and container or Ray access as applicable. Installing the Python package does not install ROCm/CUDA, profilers, or model weights.

  5. Read the repository compatibility matrix before making version claims. Treat ROCm or hardware not listed there as unverified until tested.

  6. Record the exact config, model revision, image, environment variables, GPU allocation, and Magpie commit for benchmark comparisons.

Analyze a kernel

Prefer a config when correctness or profiler settings matter:

bash
magpie analyze --kernel-config path/to/kernel.yaml

For a quick single-kernel run:

bash
magpie analyze path/to/kernel.hip --type hip --testcase "./run_test.sh"

Supported public kernel types are hip, cuda, pytorch, and triton. Use --no-perf only when the user wants correctness or execution validation without profiling.

Do not equate successful execution with numerical correctness. Supply a representative testcase whenever an optimized result will be accepted or rejected.

Compare kernel variants

Compare at least two implementations and identify the baseline explicitly:

bash
magpie compare --kernel-config path/to/compare.yaml

Keep inputs, tolerances, warmup, iteration count, GPU allocation, and profiler settings identical across candidates. Reject candidates that fail correctness before considering performance rankings.

For PyTorch without a testcase, Magpie's built-in check only verifies that each result is finite; it does not prove numerical equivalence between variants. Require a testcase for numerical validation.

Benchmark inference

Prefer a checked-in benchmark config:

bash
magpie benchmark --benchmark-config path/to/benchmark.yaml

The stable public CLI supports vllm, sglang, and atom. It supports direct docker and local run modes; use YAML configuration and the repository's Ray examples for distributed execution. Do not advertise integrations that exist only in internal enums or partial code paths as stable.

Enable profiling deliberately: profiler runs perturb latency and should not replace a clean baseline. Compare throughput, completed requests, TTFT, TPOT, ITL, and end-to-end latency using equivalent workloads.

Show full SKILL.md (432 more words)Show less

Post-process traces with TraceLens

Enable TraceLens in the profiled benchmark YAML; torch traces are its required input:

yaml
benchmark:
  profiler:
    torch_profiler:
      enabled: true
    tracelens:
      enabled: true
      analysis_mode: inference
      analysis_stages: all
      export_format: csv

Use analysis_mode: inference for vLLM/SGLang. It splits the rank-0 trace into prefilldecode, decode, and prefill stages when available, runs TraceLens post-processing, and writes full stage reports plus compact *_kernel_roofline_simple.csv files under the benchmark workspace's tracelens/ directory. For direct PyTorch trace reporting, use analysis_mode: pytorch.

Open the compact roofline CSVs first. Rank rows by kernel_time_ms_sum or time_pct; then use roofline_bound, arithmetic intensity, achieved TFLOP/s or TB/s, and pct_roofline_mean to form an optimization hypothesis. Confirm benchmark_report.json.tracelens_analysis has outputs and no error before treating post-processing as successful. Use analysis_mode: pytorch when the task specifically needs the legacy direct single-rank or multi-rank collective reports.

Magpie's integrated TraceLens stage produces CSV/Excel analysis artifacts, not an agent-written analysis.md. If the user requests a prioritized agentic report, pass the captured trace to the separate tracelens-analysis-orchestrator skill when installed; keep that result distinct from Magpie's benchmark report.

Analyze existing traces and find source

Run standalone gap analysis with --trace-dir directly on benchmark:

bash
magpie benchmark \
  --trace-dir path/to/torch_trace \
  --top-k 20 \
  --find-kernel-sources \
  --kernel-source-repos path/to/repository

Do not insert a gap-analysis positional token; it is not a CLI subcommand. Inspect the generated aggregate and per-rank CSVs, and preserve source-mapping confidence rather than assuming every normalized kernel name maps uniquely.

Drive the optimization loop

  1. Run an unprofiled baseline benchmark and save its config and report.
  2. Repeat with torch profiling and TraceLens inference post-processing enabled.
  3. Review stage-level TraceLens roofline summaries to classify dominant operations and likely compute, memory, or communication limits.
  4. Run gap analysis over the representative steady-state window to rank concrete kernels.
  5. Select bottlenecks by total contribution, not only single-dispatch duration.
  6. Map the selected kernel to source and an executable testcase.
  7. Generate isolated candidate implementations; preserve the baseline.
  8. Use analyze for iteration, then compare with correctness gates to rank candidates.
  9. Re-run the original unprofiled benchmark with the winning candidate and the same workload. Report both kernel-level and end-to-end changes, including regressions.

Stop before claiming success if correctness is unproven, the benchmark inputs changed, the source mapping is uncertain, or the end-to-end improvement is within run-to-run noise.

Use MCP tools when available

Prefer Magpie MCP tools for structured agent workflows such as hardware inspection, kernel discovery, config generation, analyze/compare, optimization suggestions, result lookup, report comparison, Ray job management, and benchmark batches.

Do not pass a CLI analyze_report.json wrapper directly to an MCP tool that expects one result object's performance_state and performance_result. Do not assume every CLI option exists in MCP; kernel-source enrichment is currently exposed by the CLI gap-analysis path.

Additional resources

© amd, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 9 other files in skills/magpie-kernel-evaluator of amd/skills.

  • SKILL.md
  • .federated.json
  • evals/evals.json
  • evals/files/analyze-simple-hip/examples/simple_hip_test/analyze_default.yaml
  • evals/files/analyze-simple-hip/examples/simple_hip_test/vector_add.hip
  • evals/files/compare-hip-variants/baseline/vector_add.hip
  • evals/files/compare-hip-variants/candidate/vector_add.hip
  • examples.md
  • reference.md
  • skill-card.md

Open the folder on GitHubat commit 6c92b41

Compare with similar skills

Magpie Kernel Evaluator next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Magpie Kernel Evaluator compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Magpie Kernel Evaluator this skillamd/skills398—~2.3kAutomated safety check: PassMIT
Graphsignalgraphsignal/graphsignal257—~6.2kAutomated safety check: PassApache-2.0
LLM Torch Profiler Trace AnalysisBBuf/AI-Infra-Auto-Driven-SKILLS911—~2.8kAutomated safety check: PassNone
LLM Serving Framework BenchmarkBBuf/AI-Infra-Auto-Driven-SKILLS911—~7.5kAutomated safety check: PassNone
LLM Pipeline Profiler AnalysisBBuf/AI-Infra-Auto-Driven-SKILLS911—~3.9kAutomated safety check: PassNone
Hyperpod Version Checkerawslabs/agent-plugins9151 repos~910Automated safety check: PassApache-2.0

Similar skills

  • Graphsignal

    graphsignal/graphsignal

    Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

    257 GitHub stars~6.2k tokensUpdated 10 days ago
    AI & LLM EngineeringAuto-check passed
  • LLM Torch Profiler Trace Analysis

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.

    911 GitHub stars~2.8k tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check passed
  • LLM Serving Framework Benchmark

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Compares SGLang, vLLM, TensorRT-LLM and TokenSpeed on one model and workload, searching server flags to find the best deployment command within a latency SLA.

    911 GitHub stars~7.5k tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check passed
  • LLM Pipeline Profiler Analysis

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Breaks LLM torch profiler traces down by forward pass, layer and kernel, with timing tables and Perfetto time ranges for the layers you want to inspect.

    911 GitHub stars~3.9k tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check passed
  • Hyperpod Version Checker

    awslabs/agent-plugins

    Official

    Check and compare software component versions on SageMaker HyperPod cluster nodes - NVIDIA drivers, CUDA toolkit, cuDNN, NCCL, EFA, AWS OFI NCCL, GDRCopy, MPI, Neuron SDK (Trainium/Inferentia)…

    915 GitHub starsUsed in 1 repo~910 tokens
    AI & LLM EngineeringAuto-check passed
  • Collect and normalize environment facts (OS, Python, GPU, CUDA/ROCm, container state) before Quark installation or PTQ planning.

    181 GitHub stars~1.4k tokensUpdated 10 days ago
    AI & LLM EngineeringAuto-check passed

More from amd/skills

All 9 skills in this repo
  • Inspects and tunes the shared-vs-dedicated memory split on AMD Ryzen APUs with unified memory (UMA) so larger LLMs and image-gen models fit on the iGPU, or so reserved GPU memory is returned to the…

    398 GitHub stars~2.6k tokensUpdated today
    Auto-check passed
  • Turns a natural-language description of routing intent into a valid Lemonade collection.router policy JSON.

    398 GitHub stars~4k tokensUpdated today
    Auto-check passed
  • Local AI Use

    amd/skills

    Makes this agent generate images, transcribe audio, and synthesize speech on the user's own machine through a local Lemonade Server instead of a paid cloud API.

    398 GitHub stars~5k tokensUpdated today
    Auto-check: notes
  • Serves AI models on AMD Instinct GPU hardware using vLLM. An agent skill from amd/skills.

    398 GitHub stars~4k tokensUpdated today
    Auto-check: notes
  • Serves an LLM on a supported AMD EPYC server CPU using vLLM with zentorch, in Docker, Podman, or conda.

    398 GitHub stars~5.7k tokensUpdated today
    Auto-check: notes
  • Autonomously optimizes end-to-end LLM inference throughput on AMD Instinct GPUs and reports a validated gain, using the Hyperloom multi-agent optimizer.

    398 GitHub stars~1.7k tokensUpdated today
    Auto-check: notes

Questions about Magpie Kernel Evaluator

What does Magpie Kernel Evaluator do?

Benchmarks LLM inference and drives GPU kernel optimization with Magpie. Magpie Kernel Evaluator is an agent skill from amd/skills. Benchmarks LLM inference and drives GPU kernel optimization with Magpie.

When should I use Magpie Kernel Evaluator?

Magpie Kernel Evaluator fits situations like: the user wants to benchmark vLLM; capture torch traces; post-process inference traces with TraceLens into prefill/decode and roofline reports; identify top bottleneck kernels.

How do I install Magpie Kernel Evaluator in Claude Code?

Run `npx skills add amd/skills --skill magpie-kernel-evaluator -a claude-code`. Or copy the skill folder (skills/magpie-kernel-evaluator in amd/skills) into .claude/skills/magpie-kernel-evaluator in your project. Claude Code loads it when a task matches its description.

How do I install Magpie Kernel Evaluator in Codex?

Run `npx skills add amd/skills --skill magpie-kernel-evaluator -a codex`. Or copy the skill folder (skills/magpie-kernel-evaluator in amd/skills) into .agents/skills/magpie-kernel-evaluator in your project. Codex loads it when a task matches its description.

Can I use Magpie Kernel Evaluator in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add amd/skills --skill magpie-kernel-evaluator -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/magpie-kernel-evaluator, .gemini/skills/magpie-kernel-evaluator, .github/skills/magpie-kernel-evaluator and .opencode/skills/magpie-kernel-evaluator in your project.

What does Magpie Kernel Evaluator need to run?

Going by SKILL.md and its folder, Magpie Kernel Evaluator needs the command-line tools its instructions call (python). Our summary lists: Python 3; Docker.

Does Magpie Kernel Evaluator access the network?

SKILL.md names 1 domain. In commands or code: github.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Magpie Kernel Evaluator safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Magpie Kernel Evaluator use?

Magpie Kernel Evaluator is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Magpie Kernel Evaluator use?

About 2.3k tokens (SKILL.md is roughly 9.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Magpie Kernel Evaluator?

Skills that share tags, products or a category with Magpie Kernel Evaluator: Graphsignal (graphsignal/graphsignal, 257 stars), LLM Torch Profiler Trace Analysis (BBuf/AI-Infra-Auto-Driven-SKILLS, 911 stars), LLM Serving Framework Benchmark (BBuf/AI-Infra-Auto-Driven-SKILLS, 911 stars) and LLM Pipeline Profiler Analysis (BBuf/AI-Infra-Auto-Driven-SKILLS, 911 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Magpie Kernel Evaluator?

amd (a GitHub organization) maintains it in amd/skills, which has 398 GitHub stars. The repository holds 9 skills in this directory. The repository was last updated on October 7, 2026.

Source: amd/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.