Agent skill

Model Deploy

by sohu-mptc in sohu-mptc/FlashRec

Deploy and serve GenRec checkpoints with FlashRec (install, serve.sh, SID trie, wide-beam knobs, health/curl, FP8, profiling).

Apache-2.0Auto-check passedAI & LLM Engineering

Install Model Deploy

skills CLI
$ npx skills add sohu-mptc/FlashRec --skill model-deploy -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install sohu-mptc/FlashRec model-deploy --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/sohu-mptc/FlashRec.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/model-deploy .claude/skills/model-deploy && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
model-deploy
GitHub stars
107
Token cost
~974 tokens
SKILL.md length
221 words
Files
1
Skills in repo
5
Repo updated
First seen
Licence
Apache-2.0

At a glance

Deploy and serve GenRec checkpoints with FlashRec (install, serve.sh, SID trie, wide-beam knobs, health/curl, FP8, profiling).

  • Works in 5 steps: python -m flashrec.check_env:CUDA /… → /health 不通:看进程是否还在、端口、CUDA graph 捕获是否卡在启动 → 非法 SID 或超长输出:确认 SID_VOCAB_FILE 对应该… → …
  • The user asks to 部署模型
  • SKILL.md covers 前置, 安装, 启动服务 and 就绪检查与请求, plus 3 more sections
  • Calls bash, python and pip

What it does

Model Deploy is an agent skill from sohu-mptc/FlashRec. Deploy and serve GenRec checkpoints with FlashRec (install, serve.sh, SID trie, wide-beam knobs, health/curl, FP8, profiling). Use when the user asks to 部署模型, 启动服务, 上线, serve, launch FlashRec, MODELPATH, /v1/chat/completions, or expose an HTTP beam-search endpoint.

Its SKILL.md is about 970 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering LLM API integration. It works with CUDA. The repository describes itself as: FlashRec is a CUDA-graph engine for generative recommendation: wide beam search (3–5 SID steps, n=50–512+) over a trie-constrained catalog, in-process FP8 serving, and ranked… The licence is Apache-2.0.

When your agent uses it

  • The user asks to 部署模型
  • Launch FlashRec
  • /v1/chat/completions
  • Expose an HTTP beam-search endpoint

Example prompts

  • “/model-deploy”

Requirements

  • Python 3
  • Docker

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. python -m flashrec.check_env:CUDA / flashinfer / sgl-kernel / 驱动
  2. /health 不通:看进程是否还在、端口、CUDA graph 捕获是否卡在启动
  3. 非法 SID 或超长输出:确认 SID_VOCAB_FILE 对应该 checkpoint;启动日志应有 Inferred --sid ...。tokenizer 命名不同时才设 SID=
  4. 宽 beam 吞吐不随并发上升:槽位不够(BATCH_SLOTS < n × 并发)
  5. 新架构加载失败:当前 ModelEngine 写死 Qwen3ForCausalLM,见 add-model

What it can do on your machine

Read from SKILL.md and the folder at commit 1089682. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • bash
    • python
    • pip
    • curl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip and curl, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Model Deploy loads about 974 tokens when it runs. Until then it costs about 70 tokens; SKILL.md has 221 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~70
When it runs · the whole SKILL.md, loaded when a task matches
~974

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from sohu-mptc/FlashRec at commit 1089682, republished under its Apache-2.0 licence (© sohu-mptc). 221 words, ~974 tokens.

Download SKILL.mdSave it as .claude/skills/model-deploy/SKILL.md (or your agent's skills folder).
name
model-deploy
description
Deploy and serve GenRec checkpoints with FlashRec (install, serve.sh, SID trie, wide-beam knobs, health/curl, FP8, profiling). Use when the user asks to 部署模型, 启动服务, 上线, serve, launch FlashRec, MODEL_PATH, /v1/chat/completions, or expose an HTTP beam-search endpoint.

部署 FlashRec

默认单进程部署。不要引入张量并行、多卡、Docker 编排或鉴权网关,除非用户明确要求。 完整旋钮见 docs/configuration.zh-CN.md。目录与评测分别走 sid-catalog、recif-eval。

前置

  • Linux + NVIDIA GPU + CUDA 12 工具链;Python ≥ 3.10
  • HuggingFace 格式 checkpoint(config.json + tokenizer + safetensors)
  • 当前只支持 Qwen3 dense(0.6B–14B、OneRec-1.7B、同架构 GenRec)。MoE / 其他架构走 add-model
  • 32 GB 显存下 FP8 大约到 14B

先确认环境:

bash
python -m flashrec.check_env

安装

bash
pip install -e .
# 或 wheel:
bash scripts/build_wheel.sh && pip install dist/flashrec-*.whl

启动服务

默认入口是 scripts/serve.sh(会设 PYTHONPATH 并展开调度环境变量)。 可调参数模板:scripts/serve.env.example(复制为 scripts/serve.env 或设 SERVE_ENV)。

服务配置是 MODEL_PATH + SID_VOCAB_FILE。引擎从 tokenizer 推断 SID 布局 (<s_a_0> codebook + <|sid_begin|> / <|sid_end|>)。不设 catalog 时做全词表 无约束解码(无 trie),只适合冒烟。

bash
SID_VOCAB_FILE=data/catalogs/sid2pid_beamrec_l4.json \
MODEL_PATH=/path/to/model bash scripts/serve.sh
# 等价:
flashrec --serve --model-path /path/to/model --port 8000 \
  --sid-vocab-file data/catalogs/sid2pid_beamrec_l4.json

常用覆盖:HOST、PORT、CUDA_VISIBLE_DEVICES、QUANTIZATION、KV_CACHE_DTYPE、 MEM_FRACTION_STATIC、CUDA_GRAPH_MAX_BS、BATCH_SLOTS、BEAM_WIDTH、 EXTRA_SERVER_ARGS。

OneRec-1.7B(文档中的服务配置)

先构建 catalog(见 sid-catalog),再:

bash
CUDA_VISIBLE_DEVICES=0 \
MODEL_PATH=/path/to/OneRec-1.7B PORT=8000 HOST=127.0.0.1 \
QUANTIZATION=fp8 KV_CACHE_DTYPE=fp8_e4m3 \
CUDA_GRAPH_MAX_BS=800 BATCH_SLOTS=800 \
SID_VOCAB_FILE=data/catalogs/sid2pid_beamrec_l4.json \
BEAM_WIDTH=50 \
  bash scripts/serve.sh

tokenizer 不用 <s_a_0> / <|sid_begin|> 约定时,再设 SID=START:END/SIZE,... 覆盖推断。

宽 beam(n = 512 / n = 1000)

默认槽位 800,一波只能进一个 512-beam 请求,且纯 LPM 可能饿死短 prompt。 n = 1000 用同一套 4096 槽位(约 4 路并发):

bash
CUDA_GRAPH_MAX_BS=4096 BATCH_SLOTS=4096 LPM_AGING_MS=150 \
  BEAM_WIDTH=512 MODEL_PATH=/path/to/model \
  SID_VOCAB_FILE=data/catalogs/sid2pid_beamrec_l4.json \
  bash scripts/serve.sh

--beam-width 决定 fused-expand 的捕获宽度;graph 尺寸会扩成 k × n。 单实例只服务一种主力 beam 宽度。

就绪检查与请求

bash
curl -sf http://127.0.0.1:8000/health
# {"status":"ok"}

curl -s http://127.0.0.1:8000/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{"messages":[{"role":"user","content":"..."}],"n":50,"max_tokens":5,"temperature":0}'
  • n:该请求 beam 宽度(覆盖服务端 --beam-width)
  • temperature=0:确定性 top-k;>0 为 Gumbel top-k,噪声只用于排序,返回的 sglext.sequence_score 不含噪声
  • 离线单次:flashrec --model-path ... --sid-vocab-file ... --prompt "..." --beam-width 50 --max-tokens 5

Profiling(与 sglang.bench_serving --profile 对齐):POST /start_profile、POST /stop_profile。 不要在 HTTP 线程上包 torch.profiler。

精度与显存

目标设置
默认生产--quantization fp8(W8A8 per-channel)+ --kv-cache-dtype fp8_e4m3
纯 BF16--quantization 传非 fp8 的值
预量化 FP8 checkpoint带 weight_scale 即可加载
OOM降 --mem-fraction-static,或改小 --cuda-graph-max-bs / --batch-slots / --beam-width

FP8 在小模型窄 beam 上不一定更快,主要省权重与 KV 显存。 OneRec-1.7B 等 BF16 训练 的 checkpoint 走加载时量化时,相对 HuggingFace 的 beam 重叠会下降,这不是框架问题;生产建议用 FP8 训练(或带 weight_scale 的预量化)权重。对照见 docs/baselines.zh-CN.md。

安全

服务无鉴权,默认绑 127.0.0.1。只在可信网或反向代理后才设 HOST=0.0.0.0。

故障排查

  1. python -m flashrec.check_env:CUDA / flashinfer / sgl-kernel / 驱动
  2. /health 不通:看进程是否还在、端口、CUDA graph 捕获是否卡在启动
  3. 非法 SID 或超长输出:确认 SID_VOCAB_FILE 对应该 checkpoint;启动日志应有 Inferred --sid ...。tokenizer 命名不同时才设 SID=
  4. 宽 beam 吞吐不随并发上升:槽位不够(BATCH_SLOTS < n × 并发)
  5. 新架构加载失败:当前 ModelEngine 写死 Qwen3ForCausalLM,见 add-model

© sohu-mptc, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/model-deploy of sohu-mptc/FlashRec.

Open the folder on GitHubat commit 1089682

Compare with similar skills

Model Deploy next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Model Deploy compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Model Deploy this skillsohu-mptc/FlashRec107—~974Automated safety check: PassApache-2.0
Azure AI Projects Python SDKmicrosoft/skills3.1k6 repos~2.8kAutomated safety check: PassMIT
Esmfold2JimLiu/science-skills2274 repos~2.5kAutomated safety check: PassApache-2.0
Dingo VerifyMigoXLab/dingo757—~833Automated safety check: PassApache-2.0
MUSA GPU Training Optimizeropen-infra-skills/infra-skills141—~1.7kAutomated safety check: PassApache-2.0
Benchmark TuneMesh-LLM/mesh-llm3.5k—~1.6kAutomated safety check: PassApache-2.0

Similar skills

  • Official

    Reference for building on Microsoft Foundry with the azure-ai-projects Python SDK: project clients, versioned agents, evaluations, connections, datasets and indexes.

    3.1k GitHub starsUsed in 6 repos~2.8k tokens
    AI & LLM EngineeringAuto-check passed
  • Esmfold2

    JimLiu/science-skills

    Biohub ESMFold2 / ESMFold2-Fast all-atom co-folding (Candido et al.

    227 GitHub starsUsed in 4 repos~2.5k tokens
    AI & LLM EngineeringAuto-check passed
  • Dingo Verify

    MigoXLab/dingo

    A skill your agent uses when the user wants to fact-check an article or verify factual claims in a document.

    757 GitHub stars~833 tokensUpdated 10 days ago
    AI & LLM EngineeringAuto-check passed
  • MUSA GPU Training Optimizer

    open-infra-skills/infra-skills

    Profiles, benchmarks and tunes AI training workloads on Moore Threads MUSA GPUs with a measurement-first process that keeps model behavior unchanged.

    141 GitHub stars~1.7k tokensUpdated 3 mo ago
    AI & LLM EngineeringAuto-check passed
  • Benchmark Tune

    Mesh-LLM/mesh-llm

    A skill your agent uses when running, debugging, interpreting, or documenting mesh-llm benchmark tune model-serving throughput trials, including choosing…

    3.5k GitHub stars~1.6k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Cuda Kernel Optimizer

    KernelFlow-ops/cuda-optimized-skill

    Iteratively optimize a CUDA/CUTLASS/Triton kernel only when strict on-device compilation, correctness, timing, and NCU evidence gates pass.

    212 GitHub stars~4.3k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed

More from sohu-mptc/FlashRec

  • Recif Eval

    sohu-mptc/FlashRec

    Run RecIF beam×concurrency eval and FlashRec vs SGLang/vLLM/TRT-LLM baselines.

    107 GitHub stars~915 tokensUpdated 7 days ago
    Auto-check passed
  • Sid Catalog

    sohu-mptc/FlashRec

    Build SID trie catalogs for FlashRec from OpenOneRec RecIF packed mappings.

    107 GitHub stars~468 tokensUpdated 7 days ago
    Auto-check passed
  • Profile Serving

    sohu-mptc/FlashRec

    Capture torch.profiler traces on a running FlashRec server. An agent skill from sohu-mptc/FlashRec.

    107 GitHub stars~312 tokensUpdated 7 days ago
    Auto-check passed
  • Add Model

    sohu-mptc/FlashRec

    给 FlashRec 引擎接入一个新模型架构(新的 HF checkpoint / 非 Qwen3 结构)。涵盖模型定义、权重合并加载、FP8 双路径、融合 kernel 接线、CUDA graph 兼容、精度校验、以及压测+trace 验证闭环。当用户要"增加/支持/接入新模型"时使用。

    107 GitHub stars~1.7k tokensUpdated 7 days ago
    Auto-check passed

Works with

Questions about Model Deploy

What does Model Deploy do?

Deploy and serve GenRec checkpoints with FlashRec (install, serve.sh, SID trie, wide-beam knobs, health/curl, FP8, profiling). Model Deploy is an agent skill from sohu-mptc/FlashRec.sh, SID trie, wide-beam knobs, health/curl, FP8, profiling).

When should I use Model Deploy?

Model Deploy fits situations like: the user asks to 部署模型; launch FlashRec; /v1/chat/completions; expose an HTTP beam-search endpoint.

How do I install Model Deploy in Claude Code?

Run `npx skills add sohu-mptc/FlashRec --skill model-deploy -a claude-code`. Or copy the skill folder (.claude/skills/model-deploy in sohu-mptc/FlashRec) into .claude/skills/model-deploy in your project. Claude Code loads it when a task matches its description.

How do I install Model Deploy in Codex?

Run `npx skills add sohu-mptc/FlashRec --skill model-deploy -a codex`. Or copy the skill folder (.claude/skills/model-deploy in sohu-mptc/FlashRec) into .agents/skills/model-deploy in your project. Codex loads it when a task matches its description.

Can I use Model Deploy in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add sohu-mptc/FlashRec --skill model-deploy -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/model-deploy, .gemini/skills/model-deploy, .github/skills/model-deploy and .opencode/skills/model-deploy in your project.

What does Model Deploy need to run?

Going by SKILL.md and its folder, Model Deploy needs the command-line tools its instructions call (bash, python, pip and curl). Our summary lists: Python 3; Docker.

Does Model Deploy access the network?

SKILL.md contains no URLs. Its commands use pip and curl, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Model Deploy safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Model Deploy use?

Model Deploy is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Model Deploy use?

About 974 tokens (SKILL.md is roughly 3.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Model Deploy?

Skills that share tags, products or a category with Model Deploy: Azure AI Projects Python SDK (microsoft/skills, 3.1k stars), Esmfold2 (JimLiu/science-skills, 227 stars), Dingo Verify (MigoXLab/dingo, 757 stars) and MUSA GPU Training Optimizer (open-infra-skills/infra-skills, 141 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Model Deploy?

sohu-mptc (a GitHub organization) maintains it in sohu-mptc/FlashRec, which has 107 GitHub stars. The repository holds 5 skills in this directory. The repository was last updated on October 1, 2026.

Source: sohu-mptc/FlashRec on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.