Agent skill

Model Inference Optimize

by majiayu000 in majiayu000/spellbook

优化实际模型推理链路,将正确性对齐、分段 profiling、显存与数据搬运、TensorRT/ONNX/PyTorch 后端、attention/kernel、FP8/compile、缓存与少步采样、质量回归、GPU 成本和服务验收串成同一实验闭环。当用户要求推理提速、降低显存或 GPU 成本、复现模型效果、定位 GPU 利用率低、优化图像/视频/扩散模型或自托管 LLM 时使用,提供…

MITAuto-check passedAI & LLM Engineering

Install Model Inference Optimize

skills CLI
$ npx skills add majiayu000/spellbook --skill model-inference-optimize -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install majiayu000/spellbook model-inference-optimize --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/majiayu000/spellbook.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/model-inference-optimize .claude/skills/model-inference-optimize && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
model-inference-optimize
GitHub stars
287
Token cost
~1.1k tokens
SKILL.md length
230 words
Files
5 (incl. references)
Skills in repo
97
Repo updated
First seen
Licence
MIT

At a glance

优化实际模型推理链路,将正确性对齐、分段 profiling、显存与数据搬运、TensorRT/ONNX/PyTorch 后端、attention/kernel、FP8/compile、缓存与少步采样、质量回归、GPU 成本和服务验收串成同一实验闭环。当用户要求推理提速、降低显存或 GPU 成本、复现模型效果、定位 GPU 利用率低、优化图像/视频/扩散模型或自托管 LLM 时使用,提供…

  • Works in 6 steps: 锁定任务和实际运行路径 → 建立可信基线 → 从 profile 选择实验 → …
  • Tasks that involve Performance optimization
  • SKILL.md covers 入口与资料, Operating Contract, 1. 锁定任务和实际运行路径 and 2. 建立可信基线, plus 5 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Model Inference Optimize is an agent skill from majiayu000/spellbook. 优化实际模型推理链路,将正确性对齐、分段 profiling、显存与数据搬运、TensorRT/ONNX/PyTorch 后端、attention/kernel、FP8/compile、缓存与少步采样、质量回归、GPU 成本和服务验收串成同一实验闭环。当用户要求推理提速、降低显存或 GPU 成本、复现模型效果、定位 GPU 利用率低、优化图像/视频/扩散模型或自托管 LLM 时使用,提供 model inference performance optimization 的脱敏案例。纯 API 调用、提示词或 token 用量优化,以及电脑卡顿,不进入本技能的模型推理实验流程。

Its SKILL.md is about 1.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including reference files (for example `agents/openai.yaml`, `references/cases.md` and `references/experiment-template.md`).

It sits in AI & LLM Engineering, covering Performance optimization, Deep learning and LLM inference and serving. It works with NVIDIA AI Platform, ONNX and PyTorch. The repository describes itself as: Cross-runtime skills for Claude Code, Codex, and multi-agent workflows. The licence is MIT.

When your agent uses it

  • Tasks that involve Performance optimization
  • Tasks that involve Deep learning
  • Tasks that involve LLM inference and serving

Example prompts

  • “/model-inference-optimize”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. 锁定任务和实际运行路径
  2. 建立可信基线
  3. 从 profile 选择实验
  4. 做可以归因的优化实验
  5. 质量、成本与真实链路共同验收
  6. 交付与续跑

What it can do on your machine

Read from SKILL.md and the folder at commit ed52af7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Model Inference Optimize loads about 1.1k tokens when it runs, and up to ~6.4k if it reads all its reference files. Until then it costs about 79 tokens; SKILL.md has 230 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~79
When it runs · the whole SKILL.md, loaded when a task matches
~1.1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~6.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from majiayu000/spellbook at commit ed52af7, republished under its MIT licence (© majiayu000). 230 words, ~1,120 tokens.

Download SKILL.mdSave it as .claude/skills/model-inference-optimize/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
model-inference-optimize
description
优化实际模型推理链路,将正确性对齐、分段 profiling、显存与数据搬运、TensorRT/ONNX/PyTorch 后端、attention/kernel、FP8/compile、缓存与少步采样、质量回归、GPU 成本和服务验收串成同一实验闭环。当用户要求推理提速、降低显存或 GPU 成本、复现模型效果、定位 GPU 利用率低、优化图像/视频/扩散模型或自托管 LLM 时使用,提供 model inference performance optimization 的脱敏案例。纯 API 调用、提示词或 token 用量优化,以及电脑卡顿,不进入本技能的模型推理实验流程。

模型推理优化系统

把明确的推理工作负载优化到可验证的结果。沿用同一任务定义、实验记录和验收标准,按瓶颈选择专项,不把所有优化技术轮流试一遍。

入口与资料

当前问题从哪里开始按需读取
模型尚未跑通、迁移后效果不对任务 → 正确性对齐专项方法 的「复现与后端迁移」
延迟高、GPU 利用率低、显存不足任务 → 基线 → profile专项方法的「测量」「显存与数据流」
换 attention/kernel、FP8 或 compile先证明热点和实际运行路径专项方法的「算子与编译」
减少 steps、换蒸馏模型或近似 decoder先定义允许的质量变化专项方法的「近似推理」
复用同类优化经验按当前瓶颈选案例脱敏案例,只读相关小节
已有优化要接入服务或证明真实收益质量回归 → 完整请求 → 服务验收专项方法的「质量评测」「服务与成本」

需要记录时复制 实验模板 到本轮产物目录,删去不适用的字段;普通 bug 修复可缩成一张对比表。

模型质量/性能对比沿用本技能。纯托管 API、Prompt、token/cache 命中或 agent 工具调用优化不进入下述 GPU 实验流程:先检查当前可用技能列表,仅在外部技能 llm-app-optimize 已安装时转交;未安装时,在当前对话中固定代表请求,记录端到端延迟、用量/费用和输出质量,一次修改一个因素,并复用应用已有的回归检查验收。

全产品竞品研究仅在外部技能 peerscope 已安装时转交;未安装时,在当前对话中明确产品对比问题,从官方资料收集功能、价格和使用限制,附来源与日期,并将未实测能力标为待验证。这两个外部技能均不是 Spellbook 安装前提,不要求为完成当前任务安装它们。

Operating Contract

  • Direct actions: 在请求范围内读代码、分析现有日志、修改本地实现并运行已授权的检查和实验。
  • Escalate before: 超出既有授权的收费运行、凭证修改、占用其他任务设备、生产变更或发布。已有授权覆盖时继续执行。
  • Evidence-backed pushback: profile 显示拟优化部分占比很低或任务规格被改变时,给出 trace/输出证据和更小的替代实验。
  • Feedback loop: 记录失败假设、质量回归和撤回原因;续跑先核对旧基线是否仍有效,再选择下一项实验。

1. 锁定任务和实际运行路径

  1. 核对目标仓库、cwd、worktree、commit、未提交改动和实际服务入口。不要把当前目录默认当成模型仓库。缺少目标模型或路径时只询问影响开工的缺口,同时整理已有证据。
  2. 从真实请求追到权重/adapter、scheduler、预处理、推理、后处理、编码和响应,记录实际加载的实现。核对参数是否被读取、provider/kernel 是否实际运行;import 成功或开关存在不算命中。
  3. 固定输入集合与哈希、输出尺寸/帧数/FPS/时长/音频、精度、seed、模型版本、GPU 型号与数量、软件栈、并发和计时边界。相同 seed 不保证跨后端或跨进程逐位一致。
  4. 写明主目标和质量约束,例如「同一 2K 视频热态请求更快,构图与颜色不退化」。区分允许数值误差、允许感知变化、允许更换模型/输出规格的范围。
  5. 使用已授权设备、时间和费用范围;不占用其他任务的 GPU、不重启现有服务或自动部署。已有授权不重复询问。没有 GPU 时完成静态定位和实验准备,性能结论留为未测。

2. 建立可信基线

  • 复用仓库 runner、数据集、profiler 和检查命令;没有可复现入口时只补当前任务需要的最小入口,不搭建通用平台。
  • 先保存真实输出并核验规格和基本质量。输入或参考未对齐时先修正确性,再优化性能。
  • 分列下载/加载、engine build/compile、首次推理、同一常驻进程热态请求和文件/API 返回。每次重启进程不算模型热态。
  • 保存逐次数据、失败/OOM、原始日志与同步方法。CUDA event 测 GPU 区间,同步后的 wall time 测实际完成时间;profiler 用于归因,无 profiler 运行用于最终成绩。
  • 资源允许时 warmup 后至少三次重复作为初步判断。收益接近波动时增加交错 A/B 或 ABBA,不挑最优一次。少量样本报原始值、中位数和范围,不称生产 p95。

3. 从 profile 选择实验

拆 loading、preprocess、conditioning、模型主体、decoder、H2D/D2H、stitch、encode、网络/排队的真实耗时和显存,优先改占比大且可改变的部分。

证据优先验证
GPU 空闲而 CPU/同步耗时高CPU island、Python crop/stitch、布局、跨 provider 同步、设备常驻
权重搬运或反复加载常驻热点组件、offload 时机、跨卡只传 conditioning、并发容量
OOM 或 reserved 很高活跃张量、KV/tile/arena/workspace、session 并发总量
attention/GEMM/卷积是热点实际 backend、shape/dtype、融合、低精度、编译、专用 kernel
decoder/输出是热点批量布局、GPU 像素转换、D2H 字节数、编码;近似 decoder 独立验收
还需显著减少计算蒸馏/少步、稀疏或跨步缓存;先界定质量变化
单请求快而服务慢/成本高队列、并发、冷启动、重试、计费占用、真实请求路径

以热点占比估算收益上限:某部分占总时长 p,即使完全删除也最多节省 p 的总时长。流水线重叠时按关键路径及完整请求测量,不直接相加阶段时间。

4. 做可以归因的优化实验

  1. 写明假设、支持它的 profile、最小改动、影响阶段、质量风险和停止条件。
  2. 保持输入、规格、计时边界与硬件一致,一次改一个因素。组合方案先消融再组合复测,不相加不同基线的节省。
  3. 局部 kernel 先用真实 shape/dtype/layout 做正确性和 microbenchmark,再测模型和最终请求。为时序状态、缓存、padding 与边界输入补相关样本。
  4. 按任务约定验收:逐位/张量误差检验实现一致性,感知/任务评测检验近似方案。生成文件不等于质量通过。
  5. 记录保留、否决或证据不足及原因。无收益、质量退化或运行不稳时撤回本次实验代码,保留证据,不影响其他人的改动。
  6. 没有可信收益、预算到限或下一步超出允许质量变化时,停止搜索并交付结果。不得为达标偷偷降分辨率、缩帧、去音频或换模型。

5. 质量、成本与真实链路共同验收

  • 图像/视频先核几何、位深、颜色、时序、音画同步和输出完整性,再算指标。针对脸、文字、纹理、运动、接缝等受影响场景留样,在优化未使用的样本上最终复核。
  • 有 HR/oracle 时使用对齐后的 fidelity/perceptual 指标与可视化;无 HR 时分开报告 source consistency、无参考分数与人工判断。单人偏好不冒充 MOS。
  • 生成模型检查任务遵循、身份、运动、时序与音频;改变 steps/权重/稀疏策略/decoder 作为有质量风险的独立候选。自托管 LLM 检查任务质量及长上下文行为。
  • 在实际入口复测完整请求与输出。局部 GPU、pipeline-return、完整 MP4、HTTP 是不同口径;仅同口径计算严格提速。
  • 同时报 wall time、GPU 数、峰值显存、成功率和计费假设。单卡更慢仍可能更省 GPU-hour,多卡更快可能更贵。区分单路延迟、每路/聚合吞吐和持续实时。
  • 授权接入时沿用项目结构,核对运行版本与命中路径,完成真实请求 smoke、输出规格及相关回归。代码或 MR 不等于上线。

6. 交付与续跑

在目标仓库已有的 runs/artifacts 位置保存一份本轮记录和原始证据;没有时使用用户产物目录,必要时落在 artifacts/model-inference-optimize/<任务>-<时间>/。不要创建多份需要同步的台账。

先给结论:改了什么、同条件前后结果、质量结论;附复现命令、输入/环境/边界、逐样本证据与失败方案。明确实际完成与未运行的检查。没有有效收益也是有效结果,不虚构测量。

续跑先读同一记录,核对 commit/权重/环境/输入是否变化,再决定是否重测基线。引用历史案例注明日期、条件与限制;需要当前软件兼容性、价格或工具能力结论时核对官方资料。

对外分享的脱敏要求

分享实验记录、案例、截图或样片前,制作独立脱敏副本。移除用户名、本机/服务器路径、会话/任务 ID、仓库和分支标识、内部链接/地址、客户或项目名、原始 prompt/人物素材、凭证,以及可反查业务的日期、精确成绩、硬件拓扑和价格。可识别模型与业务的关联时改用技术角色描述。

保留问题、假设、方法、定性结果与失败教训。需要示例数值时明确标注合成数据,不把改写后的数值称为历史实测;不把私有证据链接或身份映射表带进公开包。原始证据仅保留在用户已有的私有产物位置。

© majiayu000, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (references) in skills/model-inference-optimize of majiayu000/spellbook.

  • SKILL.md
  • agents/openai.yaml
  • references/cases.md
  • references/experiment-template.md
  • references/playbook.md

Open the folder on GitHubat commit ed52af7

Compare with similar skills

Model Inference Optimize next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Model Inference Optimize compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Model Inference Optimize this skillmajiayu000/spellbook287—~1.1kAutomated safety check: PassMIT
Graphsignalgraphsignal/graphsignal257—~6.2kAutomated safety check: PassApache-2.0
Tao Port Huggingface ModelNVIDIA/skills3.5k—~4.5kAutomated safety check: NotesApache-2.0
Model Builderqualcomm/qai-appbuilder247—~4.1kAutomated safety check: PassBSD-3-Clause
Quark Create Shapeshifter Passamd/Quark181—~2.9kAutomated safety check: PassMIT
Quark Env Preflightamd/Quark181—~1.4kAutomated safety check: PassMIT

Similar skills

  • Graphsignal

    graphsignal/graphsignal

    Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

    257 GitHub stars~6.2k tokensUpdated 11 days ago
    AI & LLM EngineeringAuto-check passed
  • Official

    Integrate a HuggingFace Computer Vision model into the NVIDIA TAO Toolkit ecosystem (tao-core config, tao-pytorch trainer, tao-deploy TensorRT pipeline).

    3.5k GitHub stars~4.5k tokensUpdated today
    AI & LLM EngineeringAuto-check: notes
  • Model Builder

    qualcomm/qai-appbuilder

    QAI ModelBuilder. An agent skill from qualcomm/qai-appbuilder.

    247 GitHub stars~4.1k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Author a new ShapeShifter graph-transformation pass for AMD Quark (ONNX or PyTorch) so it conforms to the pass framework's conventions and auto-registers.

    181 GitHub stars~2.9k tokensUpdated 11 days ago
    AI & LLM EngineeringAuto-check passed
  • Collect and normalize environment facts (OS, Python, GPU, CUDA/ROCm, container state) before Quark installation or PTQ planning.

    181 GitHub stars~1.4k tokensUpdated 11 days ago
    AI & LLM EngineeringAuto-check passed
  • Benchmarks LLM inference and drives GPU kernel optimization with Magpie.

    406 GitHub stars~2.3k tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from majiayu000/spellbook

All 97 skills in this repo
  • Skill Ecosystem Doctor

    majiayu000/spellbook

    Audits and repairs how coding-agent Skills are owned, copied and exposed across runtimes, from canonical sources to quarantine and retirement.

    287 GitHub stars~3k tokensUpdated yesterday
    Auto-check passed
  • AGENTS.md Scaffold

    majiayu000/spellbook

    Scans a repository for real evidence and proposes, or on request writes, a small stack of root and scoped AGENTS.md files with validation commands and generated-file boundaries.

    287 GitHub stars~1.5k tokensUpdated yesterday
    Auto-check passed
  • Product Demo Builder

    majiayu000/spellbook

    Plans, produces or diagnoses evidence-backed product demo videos: script, capture plan, pacing checks and verified final media built on real product behavior.

    287 GitHub stars~3.3k tokensUpdated yesterday
    Auto-check passed
  • Flowguard Task Guard

    majiayu000/spellbook

    Single entry point that routes long or ambiguous agent tasks, checks live state, bounds autonomous loops and leaves a resumable handoff.

    287 GitHub stars~2.1k tokensUpdated yesterday
    Auto-check passed
  • npm Supply Chain Check

    majiayu000/spellbook

    Scans a repository, its lockfiles and node_modules for known malicious npm package versions and install-time indicators, using a read-only Python scanner.

    287 GitHub stars~1.5k tokensUpdated yesterday
    Auto-check passed
  • Product Manager Toolkit

    majiayu000/spellbook

    Product management helpers: a RICE scoring script, an interview transcript analyzer and PRD templates for prioritizing features, synthesizing research and writing requirements.

    287 GitHub stars~2.2k tokensUpdated yesterday
    Auto-check passed

Questions about Model Inference Optimize

What does Model Inference Optimize do?

优化实际模型推理链路,将正确性对齐、分段 profiling、显存与数据搬运、TensorRT/ONNX/PyTorch 后端、attention/kernel、FP8/compile、缓存与少步采样、质量回归、GPU 成本和服务验收串成同一实验闭环。当用户要求推理提速、降低显存或 GPU 成本、复现模型效果、定位 GPU 利用率低、优化图像/视频/扩散模型或自托管 LLM 时使用,提供…. Model Inference Optimize is an agent skill from majiayu000/spellbook.

When should I use Model Inference Optimize?

Model Inference Optimize fits situations like: tasks that involve Performance optimization; tasks that involve Deep learning; tasks that involve LLM inference and serving.

How do I install Model Inference Optimize in Claude Code?

Run `npx skills add majiayu000/spellbook --skill model-inference-optimize -a claude-code`. Or copy the skill folder (skills/model-inference-optimize in majiayu000/spellbook) into .claude/skills/model-inference-optimize in your project. Claude Code loads it when a task matches its description.

How do I install Model Inference Optimize in Codex?

Run `npx skills add majiayu000/spellbook --skill model-inference-optimize -a codex`. Or copy the skill folder (skills/model-inference-optimize in majiayu000/spellbook) into .agents/skills/model-inference-optimize in your project. Codex loads it when a task matches its description.

Can I use Model Inference Optimize in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add majiayu000/spellbook --skill model-inference-optimize -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/model-inference-optimize, .gemini/skills/model-inference-optimize, .github/skills/model-inference-optimize and .opencode/skills/model-inference-optimize in your project.

What does Model Inference Optimize need to run?

SKILL.md names no scripts, command-line tools or credentials: Model Inference Optimize is instructions for the agent only. Our summary lists: Python 3.

Does Model Inference Optimize access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Model Inference Optimize safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Model Inference Optimize use?

Model Inference Optimize is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Model Inference Optimize use?

About 1.1k tokens (SKILL.md is roughly 4.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 5.3k tokens, read only when the agent opens those files.

What are the alternatives to Model Inference Optimize?

Skills that share tags, products or a category with Model Inference Optimize: Graphsignal (graphsignal/graphsignal, 257 stars), Tao Port Huggingface Model (NVIDIA/skills, 3.5k stars), Model Builder (qualcomm/qai-appbuilder, 247 stars) and Quark Create Shapeshifter Pass (amd/Quark, 181 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Model Inference Optimize?

majiayu000 (a GitHub user) maintains it in majiayu000/spellbook, which has 287 GitHub stars. The repository holds 97 skills in this directory. The repository was last updated on October 8, 2026.

Source: majiayu000/spellbook on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.