Azure AI Projects Python SDK
microsoft/skills
Reference for building on Microsoft Foundry with the azure-ai-projects Python SDK: project clients, versioned agents, evaluations, connections, datasets and indexes.
Deploy and serve GenRec checkpoints with FlashRec (install, serve.sh, SID trie, wide-beam knobs, health/curl, FP8, profiling).
$ npx skills add sohu-mptc/FlashRec --skill model-deploy -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install sohu-mptc/FlashRec model-deploy --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/sohu-mptc/FlashRec.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/model-deploy .claude/skills/model-deploy && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "model-deploy" agent skill from https://github.com/sohu-mptc/FlashRec/tree/main/.claude/skills/model-deploy into .claude/skills/model-deploy/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "model-deploy", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/sohu-mptc/FlashRec/tree/main/.claude/skills/model-deployType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add sohu-mptc/FlashRec --skill model-deploy -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install sohu-mptc/FlashRec model-deploy --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/sohu-mptc/FlashRec.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills/model-deploy .agents/skills/model-deploy && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "model-deploy" agent skill from https://github.com/sohu-mptc/FlashRec/tree/main/.claude/skills/model-deploy into .agents/skills/model-deploy/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "model-deploy", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add sohu-mptc/FlashRec --skill model-deploy -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install sohu-mptc/FlashRec model-deploy --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/sohu-mptc/FlashRec.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills/model-deploy .cursor/skills/model-deploy && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "model-deploy" agent skill from https://github.com/sohu-mptc/FlashRec/tree/main/.claude/skills/model-deploy into .cursor/skills/model-deploy/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "model-deploy", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/sohu-mptc/FlashRec.git --path .claude/skills/model-deploy--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add sohu-mptc/FlashRec --skill model-deploy -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install sohu-mptc/FlashRec model-deploy --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/sohu-mptc/FlashRec.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills/model-deploy .gemini/skills/model-deploy && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "model-deploy" agent skill from https://github.com/sohu-mptc/FlashRec/tree/main/.claude/skills/model-deploy into .gemini/skills/model-deploy/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "model-deploy", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install sohu-mptc/FlashRec model-deployInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add sohu-mptc/FlashRec --skill model-deploy -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/sohu-mptc/FlashRec.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills/model-deploy .github/skills/model-deploy && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "model-deploy" agent skill from https://github.com/sohu-mptc/FlashRec/tree/main/.claude/skills/model-deploy into .github/skills/model-deploy/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "model-deploy", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add sohu-mptc/FlashRec --skill model-deploy -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install sohu-mptc/FlashRec model-deploy --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/sohu-mptc/FlashRec.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills/model-deploy .opencode/skills/model-deploy && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "model-deploy" agent skill from https://github.com/sohu-mptc/FlashRec/tree/main/.claude/skills/model-deploy into .opencode/skills/model-deploy/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "model-deploy", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
model-deployDeploy and serve GenRec checkpoints with FlashRec (install, serve.sh, SID trie, wide-beam knobs, health/curl, FP8, profiling).
Model Deploy is an agent skill from sohu-mptc/FlashRec. Deploy and serve GenRec checkpoints with FlashRec (install, serve.sh, SID trie, wide-beam knobs, health/curl, FP8, profiling). Use when the user asks to 部署模型, 启动服务, 上线, serve, launch FlashRec, MODELPATH, /v1/chat/completions, or expose an HTTP beam-search endpoint.
Its SKILL.md is about 970 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in AI & LLM Engineering, covering LLM API integration. It works with CUDA. The repository describes itself as: FlashRec is a CUDA-graph engine for generative recommendation: wide beam search (3–5 SID steps, n=50–512+) over a trie-constrained catalog, in-process FP8 serving, and ranked… The licence is Apache-2.0.
5 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 1089682. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
bashpythonpipcurlFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use pip and curl, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Model Deploy loads about 974 tokens when it runs. Until then it costs about 70 tokens; SKILL.md has 221 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from sohu-mptc/FlashRec at commit 1089682, republished under its Apache-2.0 licence (© sohu-mptc). 221 words, ~974 tokens.
.claude/skills/model-deploy/SKILL.md (or your agent's skills folder).默认单进程部署。不要引入张量并行、多卡、Docker 编排或鉴权网关,除非用户明确要求。
完整旋钮见 docs/configuration.zh-CN.md。目录与评测分别走 sid-catalog、recif-eval。
config.json + tokenizer + safetensors)add-model先确认环境:
python -m flashrec.check_envpip install -e .
# 或 wheel:
bash scripts/build_wheel.sh && pip install dist/flashrec-*.whl默认入口是 scripts/serve.sh(会设 PYTHONPATH 并展开调度环境变量)。
可调参数模板:scripts/serve.env.example(复制为 scripts/serve.env 或设 SERVE_ENV)。
服务配置是 MODEL_PATH + SID_VOCAB_FILE。引擎从 tokenizer 推断 SID 布局
(<s_a_0> codebook + <|sid_begin|> / <|sid_end|>)。不设 catalog 时做全词表
无约束解码(无 trie),只适合冒烟。
SID_VOCAB_FILE=data/catalogs/sid2pid_beamrec_l4.json \
MODEL_PATH=/path/to/model bash scripts/serve.sh
# 等价:
flashrec --serve --model-path /path/to/model --port 8000 \
--sid-vocab-file data/catalogs/sid2pid_beamrec_l4.json常用覆盖:HOST、PORT、CUDA_VISIBLE_DEVICES、QUANTIZATION、KV_CACHE_DTYPE、
MEM_FRACTION_STATIC、CUDA_GRAPH_MAX_BS、BATCH_SLOTS、BEAM_WIDTH、
EXTRA_SERVER_ARGS。
先构建 catalog(见 sid-catalog),再:
CUDA_VISIBLE_DEVICES=0 \
MODEL_PATH=/path/to/OneRec-1.7B PORT=8000 HOST=127.0.0.1 \
QUANTIZATION=fp8 KV_CACHE_DTYPE=fp8_e4m3 \
CUDA_GRAPH_MAX_BS=800 BATCH_SLOTS=800 \
SID_VOCAB_FILE=data/catalogs/sid2pid_beamrec_l4.json \
BEAM_WIDTH=50 \
bash scripts/serve.shtokenizer 不用 <s_a_0> / <|sid_begin|> 约定时,再设
SID=START:END/SIZE,... 覆盖推断。
n = 512 / n = 1000)默认槽位 800,一波只能进一个 512-beam 请求,且纯 LPM 可能饿死短 prompt。
n = 1000 用同一套 4096 槽位(约 4 路并发):
CUDA_GRAPH_MAX_BS=4096 BATCH_SLOTS=4096 LPM_AGING_MS=150 \
BEAM_WIDTH=512 MODEL_PATH=/path/to/model \
SID_VOCAB_FILE=data/catalogs/sid2pid_beamrec_l4.json \
bash scripts/serve.sh--beam-width 决定 fused-expand 的捕获宽度;graph 尺寸会扩成 k × n。
单实例只服务一种主力 beam 宽度。
curl -sf http://127.0.0.1:8000/health
# {"status":"ok"}
curl -s http://127.0.0.1:8000/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{"messages":[{"role":"user","content":"..."}],"n":50,"max_tokens":5,"temperature":0}'n:该请求 beam 宽度(覆盖服务端 --beam-width)temperature=0:确定性 top-k;>0 为 Gumbel top-k,噪声只用于排序,返回的 sglext.sequence_score 不含噪声flashrec --model-path ... --sid-vocab-file ... --prompt "..." --beam-width 50 --max-tokens 5Profiling(与 sglang.bench_serving --profile 对齐):POST /start_profile、POST /stop_profile。
不要在 HTTP 线程上包 torch.profiler。
| 目标 | 设置 |
|---|---|
| 默认生产 | --quantization fp8(W8A8 per-channel)+ --kv-cache-dtype fp8_e4m3 |
| 纯 BF16 | --quantization 传非 fp8 的值 |
| 预量化 FP8 checkpoint | 带 weight_scale 即可加载 |
| OOM | 降 --mem-fraction-static,或改小 --cuda-graph-max-bs / --batch-slots / --beam-width |
FP8 在小模型窄 beam 上不一定更快,主要省权重与 KV 显存。
OneRec-1.7B 等 BF16 训练 的 checkpoint 走加载时量化时,相对 HuggingFace 的
beam 重叠会下降,这不是框架问题;生产建议用 FP8 训练(或带 weight_scale
的预量化)权重。对照见 docs/baselines.zh-CN.md。
服务无鉴权,默认绑 127.0.0.1。只在可信网或反向代理后才设 HOST=0.0.0.0。
python -m flashrec.check_env:CUDA / flashinfer / sgl-kernel / 驱动/health 不通:看进程是否还在、端口、CUDA graph 捕获是否卡在启动SID_VOCAB_FILE 对应该 checkpoint;启动日志应有 Inferred --sid ...。tokenizer 命名不同时才设 SID=BATCH_SLOTS < n × 并发)ModelEngine 写死 Qwen3ForCausalLM,见 add-model© sohu-mptc, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .claude/skills/model-deploy of sohu-mptc/FlashRec.
Open the folder on GitHubat commit 1089682
Model Deploy next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Model Deploy this skillsohu-mptc/FlashRec | 107 | — | ~974 | Automated safety check: Pass | Apache-2.0 | |
| Azure AI Projects Python SDKmicrosoft/skills | 3.1k | 6 repos | ~2.8k | Automated safety check: Pass | MIT | |
| Esmfold2JimLiu/science-skills | 227 | 4 repos | ~2.5k | Automated safety check: Pass | Apache-2.0 | |
| Dingo VerifyMigoXLab/dingo | 757 | — | ~833 | Automated safety check: Pass | Apache-2.0 | |
| MUSA GPU Training Optimizeropen-infra-skills/infra-skills | 141 | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | |
| Benchmark TuneMesh-LLM/mesh-llm | 3.5k | — | ~1.6k | Automated safety check: Pass | Apache-2.0 |
microsoft/skills
Reference for building on Microsoft Foundry with the azure-ai-projects Python SDK: project clients, versioned agents, evaluations, connections, datasets and indexes.
JimLiu/science-skills
Biohub ESMFold2 / ESMFold2-Fast all-atom co-folding (Candido et al.
MigoXLab/dingo
A skill your agent uses when the user wants to fact-check an article or verify factual claims in a document.
open-infra-skills/infra-skills
Profiles, benchmarks and tunes AI training workloads on Moore Threads MUSA GPUs with a measurement-first process that keeps model behavior unchanged.
Mesh-LLM/mesh-llm
A skill your agent uses when running, debugging, interpreting, or documenting mesh-llm benchmark tune model-serving throughput trials, including choosing…
KernelFlow-ops/cuda-optimized-skill
Iteratively optimize a CUDA/CUTLASS/Triton kernel only when strict on-device compilation, correctness, timing, and NCU evidence gates pass.
sohu-mptc/FlashRec
Run RecIF beam×concurrency eval and FlashRec vs SGLang/vLLM/TRT-LLM baselines.
sohu-mptc/FlashRec
Build SID trie catalogs for FlashRec from OpenOneRec RecIF packed mappings.
sohu-mptc/FlashRec
Capture torch.profiler traces on a running FlashRec server. An agent skill from sohu-mptc/FlashRec.
sohu-mptc/FlashRec
给 FlashRec 引擎接入一个新模型架构(新的 HF checkpoint / 非 Qwen3 结构)。涵盖模型定义、权重合并加载、FP8 双路径、融合 kernel 接线、CUDA graph 兼容、精度校验、以及压测+trace 验证闭环。当用户要"增加/支持/接入新模型"时使用。
Works with
Categories
Deploy and serve GenRec checkpoints with FlashRec (install, serve.sh, SID trie, wide-beam knobs, health/curl, FP8, profiling). Model Deploy is an agent skill from sohu-mptc/FlashRec.sh, SID trie, wide-beam knobs, health/curl, FP8, profiling).
Model Deploy fits situations like: the user asks to 部署模型; launch FlashRec; /v1/chat/completions; expose an HTTP beam-search endpoint.
Run `npx skills add sohu-mptc/FlashRec --skill model-deploy -a claude-code`. Or copy the skill folder (.claude/skills/model-deploy in sohu-mptc/FlashRec) into .claude/skills/model-deploy in your project. Claude Code loads it when a task matches its description.
Run `npx skills add sohu-mptc/FlashRec --skill model-deploy -a codex`. Or copy the skill folder (.claude/skills/model-deploy in sohu-mptc/FlashRec) into .agents/skills/model-deploy in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add sohu-mptc/FlashRec --skill model-deploy -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/model-deploy, .gemini/skills/model-deploy, .github/skills/model-deploy and .opencode/skills/model-deploy in your project.
Going by SKILL.md and its folder, Model Deploy needs the command-line tools its instructions call (bash, python, pip and curl). Our summary lists: Python 3; Docker.
SKILL.md contains no URLs. Its commands use pip and curl, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Model Deploy is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 974 tokens (SKILL.md is roughly 3.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Model Deploy: Azure AI Projects Python SDK (microsoft/skills, 3.1k stars), Esmfold2 (JimLiu/science-skills, 227 stars), Dingo Verify (MigoXLab/dingo, 757 stars) and MUSA GPU Training Optimizer (open-infra-skills/infra-skills, 141 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
sohu-mptc (a GitHub organization) maintains it in sohu-mptc/FlashRec, which has 107 GitHub stars. The repository holds 5 skills in this directory. The repository was last updated on October 1, 2026.
Source: sohu-mptc/FlashRec on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.