Graphsignal
graphsignal/graphsignal
Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.
优化实际模型推理链路,将正确性对齐、分段 profiling、显存与数据搬运、TensorRT/ONNX/PyTorch 后端、attention/kernel、FP8/compile、缓存与少步采样、质量回归、GPU 成本和服务验收串成同一实验闭环。当用户要求推理提速、降低显存或 GPU 成本、复现模型效果、定位 GPU 利用率低、优化图像/视频/扩散模型或自托管 LLM 时使用,提供…
$ npx skills add majiayu000/spellbook --skill model-inference-optimize -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install majiayu000/spellbook model-inference-optimize --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/majiayu000/spellbook.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/model-inference-optimize .claude/skills/model-inference-optimize && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "model-inference-optimize" agent skill from https://github.com/majiayu000/spellbook/tree/main/skills/model-inference-optimize into .claude/skills/model-inference-optimize/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "model-inference-optimize", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/majiayu000/spellbook/tree/main/skills/model-inference-optimizeType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add majiayu000/spellbook --skill model-inference-optimize -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install majiayu000/spellbook model-inference-optimize --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/majiayu000/spellbook.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/model-inference-optimize .agents/skills/model-inference-optimize && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "model-inference-optimize" agent skill from https://github.com/majiayu000/spellbook/tree/main/skills/model-inference-optimize into .agents/skills/model-inference-optimize/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "model-inference-optimize", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add majiayu000/spellbook --skill model-inference-optimize -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install majiayu000/spellbook model-inference-optimize --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/majiayu000/spellbook.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/model-inference-optimize .cursor/skills/model-inference-optimize && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "model-inference-optimize" agent skill from https://github.com/majiayu000/spellbook/tree/main/skills/model-inference-optimize into .cursor/skills/model-inference-optimize/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "model-inference-optimize", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/majiayu000/spellbook.git --path skills/model-inference-optimize--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add majiayu000/spellbook --skill model-inference-optimize -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install majiayu000/spellbook model-inference-optimize --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/majiayu000/spellbook.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/model-inference-optimize .gemini/skills/model-inference-optimize && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "model-inference-optimize" agent skill from https://github.com/majiayu000/spellbook/tree/main/skills/model-inference-optimize into .gemini/skills/model-inference-optimize/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "model-inference-optimize", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install majiayu000/spellbook model-inference-optimizeInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add majiayu000/spellbook --skill model-inference-optimize -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/majiayu000/spellbook.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/model-inference-optimize .github/skills/model-inference-optimize && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "model-inference-optimize" agent skill from https://github.com/majiayu000/spellbook/tree/main/skills/model-inference-optimize into .github/skills/model-inference-optimize/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "model-inference-optimize", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add majiayu000/spellbook --skill model-inference-optimize -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install majiayu000/spellbook model-inference-optimize --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/majiayu000/spellbook.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/model-inference-optimize .opencode/skills/model-inference-optimize && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "model-inference-optimize" agent skill from https://github.com/majiayu000/spellbook/tree/main/skills/model-inference-optimize into .opencode/skills/model-inference-optimize/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "model-inference-optimize", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
model-inference-optimize优化实际模型推理链路,将正确性对齐、分段 profiling、显存与数据搬运、TensorRT/ONNX/PyTorch 后端、attention/kernel、FP8/compile、缓存与少步采样、质量回归、GPU 成本和服务验收串成同一实验闭环。当用户要求推理提速、降低显存或 GPU 成本、复现模型效果、定位 GPU 利用率低、优化图像/视频/扩散模型或自托管 LLM 时使用,提供…
Model Inference Optimize is an agent skill from majiayu000/spellbook. 优化实际模型推理链路,将正确性对齐、分段 profiling、显存与数据搬运、TensorRT/ONNX/PyTorch 后端、attention/kernel、FP8/compile、缓存与少步采样、质量回归、GPU 成本和服务验收串成同一实验闭环。当用户要求推理提速、降低显存或 GPU 成本、复现模型效果、定位 GPU 利用率低、优化图像/视频/扩散模型或自托管 LLM 时使用,提供 model inference performance optimization 的脱敏案例。纯 API 调用、提示词或 token 用量优化,以及电脑卡顿,不进入本技能的模型推理实验流程。
Its SKILL.md is about 1.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including reference files (for example `agents/openai.yaml`, `references/cases.md` and `references/experiment-template.md`).
It sits in AI & LLM Engineering, covering Performance optimization, Deep learning and LLM inference and serving. It works with NVIDIA AI Platform, ONNX and PyTorch. The repository describes itself as: Cross-runtime skills for Claude Code, Codex, and multi-agent workflows. The licence is MIT.
6 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit ed52af7. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Model Inference Optimize loads about 1.1k tokens when it runs, and up to ~6.4k if it reads all its reference files. Until then it costs about 79 tokens; SKILL.md has 230 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from majiayu000/spellbook at commit ed52af7, republished under its MIT licence (© majiayu000). 230 words, ~1,120 tokens.
.claude/skills/model-inference-optimize/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.把明确的推理工作负载优化到可验证的结果。沿用同一任务定义、实验记录和验收标准,按瓶颈选择专项,不把所有优化技术轮流试一遍。
| 当前问题 | 从哪里开始 | 按需读取 |
|---|---|---|
| 模型尚未跑通、迁移后效果不对 | 任务 → 正确性对齐 | 专项方法 的「复现与后端迁移」 |
| 延迟高、GPU 利用率低、显存不足 | 任务 → 基线 → profile | 专项方法的「测量」「显存与数据流」 |
| 换 attention/kernel、FP8 或 compile | 先证明热点和实际运行路径 | 专项方法的「算子与编译」 |
| 减少 steps、换蒸馏模型或近似 decoder | 先定义允许的质量变化 | 专项方法的「近似推理」 |
| 复用同类优化经验 | 按当前瓶颈选案例 | 脱敏案例,只读相关小节 |
| 已有优化要接入服务或证明真实收益 | 质量回归 → 完整请求 → 服务验收 | 专项方法的「质量评测」「服务与成本」 |
需要记录时复制 实验模板 到本轮产物目录,删去不适用的字段;普通 bug 修复可缩成一张对比表。
模型质量/性能对比沿用本技能。纯托管 API、Prompt、token/cache 命中或 agent 工具调用优化不进入下述 GPU 实验流程:先检查当前可用技能列表,仅在外部技能 llm-app-optimize 已安装时转交;未安装时,在当前对话中固定代表请求,记录端到端延迟、用量/费用和输出质量,一次修改一个因素,并复用应用已有的回归检查验收。
全产品竞品研究仅在外部技能 peerscope 已安装时转交;未安装时,在当前对话中明确产品对比问题,从官方资料收集功能、价格和使用限制,附来源与日期,并将未实测能力标为待验证。这两个外部技能均不是 Spellbook 安装前提,不要求为完成当前任务安装它们。
拆 loading、preprocess、conditioning、模型主体、decoder、H2D/D2H、stitch、encode、网络/排队的真实耗时和显存,优先改占比大且可改变的部分。
| 证据 | 优先验证 |
|---|---|
| GPU 空闲而 CPU/同步耗时高 | CPU island、Python crop/stitch、布局、跨 provider 同步、设备常驻 |
| 权重搬运或反复加载 | 常驻热点组件、offload 时机、跨卡只传 conditioning、并发容量 |
| OOM 或 reserved 很高 | 活跃张量、KV/tile/arena/workspace、session 并发总量 |
| attention/GEMM/卷积是热点 | 实际 backend、shape/dtype、融合、低精度、编译、专用 kernel |
| decoder/输出是热点 | 批量布局、GPU 像素转换、D2H 字节数、编码;近似 decoder 独立验收 |
| 还需显著减少计算 | 蒸馏/少步、稀疏或跨步缓存;先界定质量变化 |
| 单请求快而服务慢/成本高 | 队列、并发、冷启动、重试、计费占用、真实请求路径 |
以热点占比估算收益上限:某部分占总时长 p,即使完全删除也最多节省 p 的总时长。流水线重叠时按关键路径及完整请求测量,不直接相加阶段时间。
在目标仓库已有的 runs/artifacts 位置保存一份本轮记录和原始证据;没有时使用用户产物目录,必要时落在 artifacts/model-inference-optimize/<任务>-<时间>/。不要创建多份需要同步的台账。
先给结论:改了什么、同条件前后结果、质量结论;附复现命令、输入/环境/边界、逐样本证据与失败方案。明确实际完成与未运行的检查。没有有效收益也是有效结果,不虚构测量。
续跑先读同一记录,核对 commit/权重/环境/输入是否变化,再决定是否重测基线。引用历史案例注明日期、条件与限制;需要当前软件兼容性、价格或工具能力结论时核对官方资料。
分享实验记录、案例、截图或样片前,制作独立脱敏副本。移除用户名、本机/服务器路径、会话/任务 ID、仓库和分支标识、内部链接/地址、客户或项目名、原始 prompt/人物素材、凭证,以及可反查业务的日期、精确成绩、硬件拓扑和价格。可识别模型与业务的关联时改用技术角色描述。
保留问题、假设、方法、定性结果与失败教训。需要示例数值时明确标注合成数据,不把改写后的数值称为历史实测;不把私有证据链接或身份映射表带进公开包。原始证据仅保留在用户已有的私有产物位置。
© majiayu000, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 4 other files (references) in skills/model-inference-optimize of majiayu000/spellbook.
Open the folder on GitHubat commit ed52af7
Model Inference Optimize next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Model Inference Optimize this skillmajiayu000/spellbook | 287 | — | ~1.1k | Automated safety check: Pass | MIT | |
| Graphsignalgraphsignal/graphsignal | 257 | — | ~6.2k | Automated safety check: Pass | Apache-2.0 | |
| Tao Port Huggingface ModelNVIDIA/skills | 3.5k | — | ~4.5k | Automated safety check: Notes | Apache-2.0 | |
| Model Builderqualcomm/qai-appbuilder | 247 | — | ~4.1k | Automated safety check: Pass | BSD-3-Clause | |
| Quark Create Shapeshifter Passamd/Quark | 181 | — | ~2.9k | Automated safety check: Pass | MIT | |
| Quark Env Preflightamd/Quark | 181 | — | ~1.4k | Automated safety check: Pass | MIT |
graphsignal/graphsignal
Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.
NVIDIA/skills
Integrate a HuggingFace Computer Vision model into the NVIDIA TAO Toolkit ecosystem (tao-core config, tao-pytorch trainer, tao-deploy TensorRT pipeline).
qualcomm/qai-appbuilder
QAI ModelBuilder. An agent skill from qualcomm/qai-appbuilder.
amd/Quark
Author a new ShapeShifter graph-transformation pass for AMD Quark (ONNX or PyTorch) so it conforms to the pass framework's conventions and auto-registers.
amd/Quark
Collect and normalize environment facts (OS, Python, GPU, CUDA/ROCm, container state) before Quark installation or PTQ planning.
amd/skills
Benchmarks LLM inference and drives GPU kernel optimization with Magpie.
majiayu000/spellbook
Audits and repairs how coding-agent Skills are owned, copied and exposed across runtimes, from canonical sources to quarantine and retirement.
majiayu000/spellbook
Scans a repository for real evidence and proposes, or on request writes, a small stack of root and scoped AGENTS.md files with validation commands and generated-file boundaries.
majiayu000/spellbook
Plans, produces or diagnoses evidence-backed product demo videos: script, capture plan, pacing checks and verified final media built on real product behavior.
majiayu000/spellbook
Single entry point that routes long or ambiguous agent tasks, checks live state, bounds autonomous loops and leaves a resumable handoff.
majiayu000/spellbook
Scans a repository, its lockfiles and node_modules for known malicious npm package versions and install-time indicators, using a read-only Python scanner.
majiayu000/spellbook
Product management helpers: a RICE scoring script, an interview transcript analyzer and PRD templates for prioritizing features, synthesizing research and writing requirements.
Works with
Categories
优化实际模型推理链路,将正确性对齐、分段 profiling、显存与数据搬运、TensorRT/ONNX/PyTorch 后端、attention/kernel、FP8/compile、缓存与少步采样、质量回归、GPU 成本和服务验收串成同一实验闭环。当用户要求推理提速、降低显存或 GPU 成本、复现模型效果、定位 GPU 利用率低、优化图像/视频/扩散模型或自托管 LLM 时使用,提供…. Model Inference Optimize is an agent skill from majiayu000/spellbook.
Model Inference Optimize fits situations like: tasks that involve Performance optimization; tasks that involve Deep learning; tasks that involve LLM inference and serving.
Run `npx skills add majiayu000/spellbook --skill model-inference-optimize -a claude-code`. Or copy the skill folder (skills/model-inference-optimize in majiayu000/spellbook) into .claude/skills/model-inference-optimize in your project. Claude Code loads it when a task matches its description.
Run `npx skills add majiayu000/spellbook --skill model-inference-optimize -a codex`. Or copy the skill folder (skills/model-inference-optimize in majiayu000/spellbook) into .agents/skills/model-inference-optimize in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add majiayu000/spellbook --skill model-inference-optimize -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/model-inference-optimize, .gemini/skills/model-inference-optimize, .github/skills/model-inference-optimize and .opencode/skills/model-inference-optimize in your project.
SKILL.md names no scripts, command-line tools or credentials: Model Inference Optimize is instructions for the agent only. Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Model Inference Optimize is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.1k tokens (SKILL.md is roughly 4.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 5.3k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Model Inference Optimize: Graphsignal (graphsignal/graphsignal, 257 stars), Tao Port Huggingface Model (NVIDIA/skills, 3.5k stars), Model Builder (qualcomm/qai-appbuilder, 247 stars) and Quark Create Shapeshifter Pass (amd/Quark, 181 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
majiayu000 (a GitHub user) maintains it in majiayu000/spellbook, which has 287 GitHub stars. The repository holds 97 skills in this directory. The repository was last updated on October 8, 2026.
Source: majiayu000/spellbook on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.