Graphsignal
graphsignal/graphsignal
Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.
Benchmarks LLM inference and drives GPU kernel optimization with Magpie.
$ npx skills add amd/skills --skill magpie-kernel-evaluator -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install amd/skills magpie-kernel-evaluator --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/amd/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/magpie-kernel-evaluator .claude/skills/magpie-kernel-evaluator && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "magpie-kernel-evaluator" agent skill from https://github.com/amd/skills/tree/main/skills/magpie-kernel-evaluator into .claude/skills/magpie-kernel-evaluator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "magpie-kernel-evaluator", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/amd/skills/tree/main/skills/magpie-kernel-evaluatorType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add amd/skills --skill magpie-kernel-evaluator -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install amd/skills magpie-kernel-evaluator --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/amd/skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/magpie-kernel-evaluator .agents/skills/magpie-kernel-evaluator && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "magpie-kernel-evaluator" agent skill from https://github.com/amd/skills/tree/main/skills/magpie-kernel-evaluator into .agents/skills/magpie-kernel-evaluator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "magpie-kernel-evaluator", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add amd/skills --skill magpie-kernel-evaluator -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install amd/skills magpie-kernel-evaluator --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/amd/skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/magpie-kernel-evaluator .cursor/skills/magpie-kernel-evaluator && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "magpie-kernel-evaluator" agent skill from https://github.com/amd/skills/tree/main/skills/magpie-kernel-evaluator into .cursor/skills/magpie-kernel-evaluator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "magpie-kernel-evaluator", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/amd/skills.git --path skills/magpie-kernel-evaluator--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add amd/skills --skill magpie-kernel-evaluator -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install amd/skills magpie-kernel-evaluator --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/amd/skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/magpie-kernel-evaluator .gemini/skills/magpie-kernel-evaluator && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "magpie-kernel-evaluator" agent skill from https://github.com/amd/skills/tree/main/skills/magpie-kernel-evaluator into .gemini/skills/magpie-kernel-evaluator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "magpie-kernel-evaluator", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install amd/skills magpie-kernel-evaluatorInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add amd/skills --skill magpie-kernel-evaluator -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/amd/skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/magpie-kernel-evaluator .github/skills/magpie-kernel-evaluator && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "magpie-kernel-evaluator" agent skill from https://github.com/amd/skills/tree/main/skills/magpie-kernel-evaluator into .github/skills/magpie-kernel-evaluator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "magpie-kernel-evaluator", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add amd/skills --skill magpie-kernel-evaluator -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install amd/skills magpie-kernel-evaluator --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/amd/skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/magpie-kernel-evaluator .opencode/skills/magpie-kernel-evaluator && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "magpie-kernel-evaluator" agent skill from https://github.com/amd/skills/tree/main/skills/magpie-kernel-evaluator into .opencode/skills/magpie-kernel-evaluator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "magpie-kernel-evaluator", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
magpie-kernel-evaluatorBenchmarks LLM inference and drives GPU kernel optimization with Magpie.
Magpie Kernel Evaluator is an agent skill from amd/skills. Benchmarks LLM inference and drives GPU kernel optimization with Magpie. Use when the user wants to benchmark vLLM, SGLang, or Atom; capture torch traces; post-process inference traces with TraceLens into prefill/decode and roofline reports; identify top bottleneck kernels or map profiler names to source; analyze or compare HIP, CUDA, PyTorch, or Triton kernels; validate and rank optimized variants; run local, container, or Ray workloads; or mentions Magpie, TraceLens, gap analysis, TTFT, TPOT, kernel evaluation…
Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 17 other files (for example `.federated.json`, `evals/evals.json` and `evals/files/analyze-simple-hip/examples/simple_hip_test/analyze_default.yaml`).
It sits in AI & LLM Engineering, covering GPU and accelerator computing, LLM inference and serving and Performance optimization. It works with SGLang, vLLM, CUDA and PyTorch. The repository describes itself as: Official AMD catalog of AI agent skills. Empower your AI agents with AMD's optimized SW stack. The licence is MIT.
3 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 6c92b41. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pythonFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
github.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Magpie Kernel Evaluator loads about 2.3k tokens when it runs. Until then it costs about 142 tokens; SKILL.md has 1,046 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from amd/skills at commit 6c92b41, republished under its MIT licence (© amd). 1,046 words, ~2,333 tokens.
.claude/skills/magpie-kernel-evaluator/SKILL.md (or your agent's skills folder). This skill also uses 9 other files; get the full folder from GitHub.Use Magpie for three connected jobs:
Describe only capabilities supported by the checked-out Magpie version. Do not infer support for an unverified ROCm, GPU, framework, or experimental integration.
| User goal | Workflow |
|---|---|
| Evaluate one implementation | analyze |
| Rank two or more implementations | compare |
| Measure model-serving performance | benchmark |
| Find expensive kernels in existing traces | standalone gap analysis |
| Explain a profiled inference workload | benchmark → TraceLens post-processing → stage/roofline review |
| Optimize an end-to-end workload | benchmark → TraceLens/gap analysis → source mapping → analyze/compare → re-benchmark |
Use a YAML config for reproducible or multi-step work. Use inline CLI arguments for small exploratory runs.
Installing this skill does not install the Magpie application (magpie-eval). For a preparation-only request, read the supplied fixtures and reference.md, write the requested plan or configuration, and identify commands that remain unverified. Do not install packages or execute a workload when the user forbids it; a missing Magpie installation does not prevent preparing a plan.
Before executing a Magpie workload:
Check magpie --help in the intended Python environment (Python 3.10+). If the command is missing, try python -m Magpie --help in that same environment; a working module entry point can be used instead of the CLI. No module named Magpie means the application is not installed in that environment. For errors inside an installed Magpie package, diagnose the reported import or dependency failure instead.
If Magpie is not installed, install it in the intended environment before continuing:
python -m pip install git+https://github.com/AMD-AGI/Magpie.gitFor an existing Magpie source checkout, use python -m pip install -e /path/to/Magpie instead. The installed skill folder and a testcase workspace are not Magpie source checkouts. Re-run magpie --help or python -m Magpie --help to verify the installation; if it fails, report the error before attempting a workload.
Inspect the installed interface for the selected workflow (substitute python -m Magpie if using the module entry point):
magpie --help
magpie analyze --help
magpie compare --help
magpie benchmark --help
magpie --gpu-infoCheck required tools, model access, GPU visibility, writable output space, and container or Ray access as applicable. Installing the Python package does not install ROCm/CUDA, profilers, or model weights.
Read the repository compatibility matrix before making version claims. Treat ROCm or hardware not listed there as unverified until tested.
Record the exact config, model revision, image, environment variables, GPU allocation, and Magpie commit for benchmark comparisons.
Prefer a config when correctness or profiler settings matter:
magpie analyze --kernel-config path/to/kernel.yamlFor a quick single-kernel run:
magpie analyze path/to/kernel.hip --type hip --testcase "./run_test.sh"Supported public kernel types are hip, cuda, pytorch, and triton. Use --no-perf only when the user wants correctness or execution validation without profiling.
Do not equate successful execution with numerical correctness. Supply a representative testcase whenever an optimized result will be accepted or rejected.
Compare at least two implementations and identify the baseline explicitly:
magpie compare --kernel-config path/to/compare.yamlKeep inputs, tolerances, warmup, iteration count, GPU allocation, and profiler settings identical across candidates. Reject candidates that fail correctness before considering performance rankings.
For PyTorch without a testcase, Magpie's built-in check only verifies that each result is finite; it does not prove numerical equivalence between variants. Require a testcase for numerical validation.
Prefer a checked-in benchmark config:
magpie benchmark --benchmark-config path/to/benchmark.yamlThe stable public CLI supports vllm, sglang, and atom. It supports direct docker and local run modes; use YAML configuration and the repository's Ray examples for distributed execution. Do not advertise integrations that exist only in internal enums or partial code paths as stable.
Enable profiling deliberately: profiler runs perturb latency and should not replace a clean baseline. Compare throughput, completed requests, TTFT, TPOT, ITL, and end-to-end latency using equivalent workloads.
Enable TraceLens in the profiled benchmark YAML; torch traces are its required input:
benchmark:
profiler:
torch_profiler:
enabled: true
tracelens:
enabled: true
analysis_mode: inference
analysis_stages: all
export_format: csvUse analysis_mode: inference for vLLM/SGLang. It splits the rank-0 trace into prefilldecode, decode, and prefill stages when available, runs TraceLens post-processing, and writes full stage reports plus compact *_kernel_roofline_simple.csv files under the benchmark workspace's tracelens/ directory. For direct PyTorch trace reporting, use analysis_mode: pytorch.
Open the compact roofline CSVs first. Rank rows by kernel_time_ms_sum or time_pct; then use roofline_bound, arithmetic intensity, achieved TFLOP/s or TB/s, and pct_roofline_mean to form an optimization hypothesis. Confirm benchmark_report.json.tracelens_analysis has outputs and no error before treating post-processing as successful. Use analysis_mode: pytorch when the task specifically needs the legacy direct single-rank or multi-rank collective reports.
Magpie's integrated TraceLens stage produces CSV/Excel analysis artifacts, not an agent-written analysis.md. If the user requests a prioritized agentic report, pass the captured trace to the separate tracelens-analysis-orchestrator skill when installed; keep that result distinct from Magpie's benchmark report.
Run standalone gap analysis with --trace-dir directly on benchmark:
magpie benchmark \
--trace-dir path/to/torch_trace \
--top-k 20 \
--find-kernel-sources \
--kernel-source-repos path/to/repositoryDo not insert a gap-analysis positional token; it is not a CLI subcommand. Inspect the generated aggregate and per-rank CSVs, and preserve source-mapping confidence rather than assuming every normalized kernel name maps uniquely.
analyze for iteration, then compare with correctness gates to rank candidates.Stop before claiming success if correctness is unproven, the benchmark inputs changed, the source mapping is uncertain, or the end-to-end improvement is within run-to-run noise.
Prefer Magpie MCP tools for structured agent workflows such as hardware inspection, kernel discovery, config generation, analyze/compare, optimization suggestions, result lookup, report comparison, Ray job management, and benchmark batches.
Do not pass a CLI analyze_report.json wrapper directly to an MCP tool that expects one result object's performance_state and performance_result. Do not assume every CLI option exists in MCP; kernel-source enrichment is currently exposed by the CLI gap-analysis path.
© amd, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 9 other files in skills/magpie-kernel-evaluator of amd/skills.
Open the folder on GitHubat commit 6c92b41
Magpie Kernel Evaluator next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Magpie Kernel Evaluator this skillamd/skills | 398 | — | ~2.3k | Automated safety check: Pass | MIT | |
| Graphsignalgraphsignal/graphsignal | 257 | — | ~6.2k | Automated safety check: Pass | Apache-2.0 | |
| LLM Torch Profiler Trace AnalysisBBuf/AI-Infra-Auto-Driven-SKILLS | 911 | — | ~2.8k | Automated safety check: Pass | None | |
| LLM Serving Framework BenchmarkBBuf/AI-Infra-Auto-Driven-SKILLS | 911 | — | ~7.5k | Automated safety check: Pass | None | |
| LLM Pipeline Profiler AnalysisBBuf/AI-Infra-Auto-Driven-SKILLS | 911 | — | ~3.9k | Automated safety check: Pass | None | |
| Hyperpod Version Checkerawslabs/agent-plugins | 915 | 1 repos | ~910 | Automated safety check: Pass | Apache-2.0 |
graphsignal/graphsignal
Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.
BBuf/AI-Infra-Auto-Driven-SKILLS
Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.
BBuf/AI-Infra-Auto-Driven-SKILLS
Compares SGLang, vLLM, TensorRT-LLM and TokenSpeed on one model and workload, searching server flags to find the best deployment command within a latency SLA.
BBuf/AI-Infra-Auto-Driven-SKILLS
Breaks LLM torch profiler traces down by forward pass, layer and kernel, with timing tables and Perfetto time ranges for the layers you want to inspect.
awslabs/agent-plugins
Check and compare software component versions on SageMaker HyperPod cluster nodes - NVIDIA drivers, CUDA toolkit, cuDNN, NCCL, EFA, AWS OFI NCCL, GDRCopy, MPI, Neuron SDK (Trainium/Inferentia)…
amd/Quark
Collect and normalize environment facts (OS, Python, GPU, CUDA/ROCm, container state) before Quark installation or PTQ planning.
amd/skills
Inspects and tunes the shared-vs-dedicated memory split on AMD Ryzen APUs with unified memory (UMA) so larger LLMs and image-gen models fit on the iGPU, or so reserved GPU memory is returned to the…
amd/skills
Turns a natural-language description of routing intent into a valid Lemonade collection.router policy JSON.
amd/skills
Makes this agent generate images, transcribe audio, and synthesize speech on the user's own machine through a local Lemonade Server instead of a paid cloud API.
amd/skills
Serves AI models on AMD Instinct GPU hardware using vLLM. An agent skill from amd/skills.
amd/skills
Serves an LLM on a supported AMD EPYC server CPU using vLLM with zentorch, in Docker, Podman, or conda.
amd/skills
Autonomously optimizes end-to-end LLM inference throughput on AMD Instinct GPUs and reports a validated gain, using the Hyperloom multi-agent optimizer.
Categories
Benchmarks LLM inference and drives GPU kernel optimization with Magpie. Magpie Kernel Evaluator is an agent skill from amd/skills. Benchmarks LLM inference and drives GPU kernel optimization with Magpie.
Magpie Kernel Evaluator fits situations like: the user wants to benchmark vLLM; capture torch traces; post-process inference traces with TraceLens into prefill/decode and roofline reports; identify top bottleneck kernels.
Run `npx skills add amd/skills --skill magpie-kernel-evaluator -a claude-code`. Or copy the skill folder (skills/magpie-kernel-evaluator in amd/skills) into .claude/skills/magpie-kernel-evaluator in your project. Claude Code loads it when a task matches its description.
Run `npx skills add amd/skills --skill magpie-kernel-evaluator -a codex`. Or copy the skill folder (skills/magpie-kernel-evaluator in amd/skills) into .agents/skills/magpie-kernel-evaluator in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add amd/skills --skill magpie-kernel-evaluator -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/magpie-kernel-evaluator, .gemini/skills/magpie-kernel-evaluator, .github/skills/magpie-kernel-evaluator and .opencode/skills/magpie-kernel-evaluator in your project.
Going by SKILL.md and its folder, Magpie Kernel Evaluator needs the command-line tools its instructions call (python). Our summary lists: Python 3; Docker.
SKILL.md names 1 domain. In commands or code: github.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Magpie Kernel Evaluator is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.3k tokens (SKILL.md is roughly 9.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Magpie Kernel Evaluator: Graphsignal (graphsignal/graphsignal, 257 stars), LLM Torch Profiler Trace Analysis (BBuf/AI-Infra-Auto-Driven-SKILLS, 911 stars), LLM Serving Framework Benchmark (BBuf/AI-Infra-Auto-Driven-SKILLS, 911 stars) and LLM Pipeline Profiler Analysis (BBuf/AI-Infra-Auto-Driven-SKILLS, 911 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
amd (a GitHub organization) maintains it in amd/skills, which has 398 GitHub stars. The repository holds 9 skills in this directory. The repository was last updated on October 7, 2026.
Source: amd/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.