SageMaker Serving Image Selection
huggingface/skills
Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.
Run RecIF beam×concurrency eval and FlashRec vs SGLang/vLLM/TRT-LLM baselines.
$ npx skills add sohu-mptc/FlashRec --skill recif-eval -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install sohu-mptc/FlashRec recif-eval --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/sohu-mptc/FlashRec.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/recif-eval .claude/skills/recif-eval && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "recif-eval" agent skill from https://github.com/sohu-mptc/FlashRec/tree/main/.claude/skills/recif-eval into .claude/skills/recif-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "recif-eval", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/sohu-mptc/FlashRec/tree/main/.claude/skills/recif-evalType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add sohu-mptc/FlashRec --skill recif-eval -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install sohu-mptc/FlashRec recif-eval --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/sohu-mptc/FlashRec.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills/recif-eval .agents/skills/recif-eval && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "recif-eval" agent skill from https://github.com/sohu-mptc/FlashRec/tree/main/.claude/skills/recif-eval into .agents/skills/recif-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "recif-eval", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add sohu-mptc/FlashRec --skill recif-eval -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install sohu-mptc/FlashRec recif-eval --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/sohu-mptc/FlashRec.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills/recif-eval .cursor/skills/recif-eval && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "recif-eval" agent skill from https://github.com/sohu-mptc/FlashRec/tree/main/.claude/skills/recif-eval into .cursor/skills/recif-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "recif-eval", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/sohu-mptc/FlashRec.git --path .claude/skills/recif-eval--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add sohu-mptc/FlashRec --skill recif-eval -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install sohu-mptc/FlashRec recif-eval --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/sohu-mptc/FlashRec.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills/recif-eval .gemini/skills/recif-eval && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "recif-eval" agent skill from https://github.com/sohu-mptc/FlashRec/tree/main/.claude/skills/recif-eval into .gemini/skills/recif-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "recif-eval", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install sohu-mptc/FlashRec recif-evalInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add sohu-mptc/FlashRec --skill recif-eval -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/sohu-mptc/FlashRec.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills/recif-eval .github/skills/recif-eval && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "recif-eval" agent skill from https://github.com/sohu-mptc/FlashRec/tree/main/.claude/skills/recif-eval into .github/skills/recif-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "recif-eval", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add sohu-mptc/FlashRec --skill recif-eval -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install sohu-mptc/FlashRec recif-eval --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/sohu-mptc/FlashRec.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills/recif-eval .opencode/skills/recif-eval && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "recif-eval" agent skill from https://github.com/sohu-mptc/FlashRec/tree/main/.claude/skills/recif-eval into .opencode/skills/recif-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "recif-eval", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
recif-evalRun RecIF beam×concurrency eval and FlashRec vs SGLang/vLLM/TRT-LLM baselines.
Recif Eval is an agent skill from sohu-mptc/FlashRec. Run RecIF beam×concurrency eval and FlashRec vs SGLang/vLLM/TRT-LLM baselines. Use when the user asks to 评测, 对照, 压测, matrix, RecIF, recall@32, invalidrate, runsglangflashrecmatrix, benchsglangcompare, or evalbeammatrix.
Its SKILL.md is about 920 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in AI & LLM Engineering, covering LLM inference and serving. It works with SGLang and vLLM. The repository describes itself as: FlashRec is a CUDA-graph engine for generative recommendation: wide beam search (3–5 SID steps, n=50–512+) over a trie-constrained catalog, in-process FP8 serving, and ranked… The licence is Apache-2.0.
Read from SKILL.md and the folder at commit 1089682. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pythonbashFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
github.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Recif Eval loads about 915 tokens when it runs. Until then it costs about 60 tokens; SKILL.md has 170 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from sohu-mptc/FlashRec at commit 1089682, republished under its Apache-2.0 licence (© sohu-mptc). 170 words, ~915 tokens.
.claude/skills/recif-eval/SKILL.md (or your agent's skills folder).公开数字用 OneRec-1.7B(快手 OneRec-1.7B × OpenOneRec RecIF-Bench video)。
对照协议与表格见 docs/baselines.zh-CN.md。
不要把不同 workload 的 QPS 直接相除。 OneRec-1.7B prompt 约 2.5k token(~500 条历史 SID);
同一引擎 n=50 并发 8 时约 26 QPS,饱和约 28 QPS。生产短 prompt(约 300 token)
n=50 饱和约 220 QPS(1.7B)/ 303 QPS(0.6B);n=1000 为 24.6 QPS(1.7B)/
32 QPS(0.6B)。见 docs/baselines.zh-CN.md。
SGLang-master 与 SGLang 0801 分列。 SGLang-master 是 PR #31626;0801 是
cswuyg/sglang feature/beam_search_update_0801。写结论、表格、加速比时不要把
两家 QPS 合成一行。
不公平对照警告: FlashRec 默认 SID trie;SGLang-master / SGLang 0801 / vLLM / HF /
TensorRT-LLM 在本仓库文档中是开放词表 beam。比吞吐时对齐 beam_width、
max_tokens、模型、硬件;比质量时必须报 invalid_rate。不要引用 FlashRec
conc=1 的 recall@32(unique-beam 塌缩)。
需要两张卡(默认 FLASHREC_GPU=0、SGL_GPU=1),且 $PYTHON 能 import 带
PR #31626 的 SGLang-master。
# 1. catalog(一次)
DATA_DIR=/path/to/OpenOneRec-RecIF/benchmark_data \
bash scripts/build_catalog.sh
# 2. 速度对照(推荐):beam {50,128,512} × 并发 {1,8,16,32}
# SGLang-master 默认 Docker 镜像 nightly-dev-cu13-20260827-20621aa1,请求走 /generate + beam_width
MODEL_PATH=/path/to/OneRec-1.7B \
DATA_DIR=/path/to/OpenOneRec-RecIF/benchmark_data \
SMOKE=1 bash scripts/bench_sglang_compare.sh
MODEL_PATH=/path/to/OneRec-1.7B \
DATA_DIR=/path/to/OpenOneRec-RecIF/benchmark_data \
bash scripts/bench_sglang_compare.sh产物:results/onerec_beam_conc_bench_<stamp>/(MATRIX_REPORT.md、
SPEED_COMPARE.md、matrix_summary.json、每格 summary.json /
candidates.csv / per_sample_metrics.csv / latencies.jsonl)。
results/ 已 gitignore。从一次 run 重生成对照表:
python scripts/summarize_sglang_compare.py results/onerec_beam_conc_bench_<stamp>可覆盖:BEAMS、CONCS、SAMPLE_SIZE、FLASHREC_GPU、SGL_GPU、FLASHREC_PORT、
SGL_PORT、PYTHON、SGLANG_LAUNCH、DOCKER、OUTDIR。FlashRec 按 n 自动设
graph / 槽位(50→800、128→2048、≥512→4096),换宽度会重启两侧服务。
FlashRec 侧只需要 SID_VOCAB_FILE(脚本默认 data/catalogs/sid2pid_beamrec_l4.json);
布局从 tokenizer 推断,不必设 SID。SGLang-master 是开放词表,客户端必须打
POST /generate + sampling_params.beam_width(eval_beam_matrix.py --engine sglang)。
服务已在跑时:
python scripts/eval_beam_matrix.py \
--engine flashrec \
--server-url http://127.0.0.1:8000 \
--data-dir /path/to/benchmark_data \
--catalog data/catalogs/sid2pid_beamrec_l4.json \
--out-dir /tmp/cell --task video --n 50 --concurrency 8 --sample-size 200--engine sglang 走 POST /generate(还要 --model-path)。其它引擎走
/v1/chat/completions。看 summary.json 的 qps、invalid_rate、
metrics.recall@32 / ndcg@32。
启动命令在 docs/baselines.zh-CN.md 文末,镜像不随仓库提供。要点:
sampling_params.beam_width(POST /generate)、
--disable-overlap-schedule、--max-running-requests ≥ (n+1) × 并发
(expand 会复制 request-pool 槽位)。速度对照用 scripts/bench_sglang_compare.shuse_beam_search=true;--max-logprobs ≥ 2n(n=32 至少 100)--max_beam_width 与请求 best_of 一致;32 GB 上 n=512 通常不支持PYTHONPATH=python python -m unittest discover -s tests -v
SGLANG_BEAM_URL=http://127.0.0.1:PORT FLASHREC_URL=http://127.0.0.1:8000 \
PYTHONPATH=python python -m unittest tests.test_parity -v
FLASHREC_DIFF_MODEL=/path/to/model \
PYTHONPATH=python python -m unittest tests.test_beam_search_diff -v© sohu-mptc, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .claude/skills/recif-eval of sohu-mptc/FlashRec.
Open the folder on GitHubat commit 1089682
Recif Eval next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Recif Eval this skillsohu-mptc/FlashRec | 107 | — | ~915 | Automated safety check: Pass | Apache-2.0 | |
| SageMaker Serving Image Selectionhuggingface/skills | 11k | 1 repos | ~4.6k | Automated safety check: Pass | Apache-2.0 | |
| Dstack Prototypingdstackai/dstack | 2.3k | — | ~1.6k | Automated safety check: Pass | MPL-2.0 | |
| Debug InferenceNVIDIA/OpenShell | 16k | — | ~1.9k | Automated safety check: Pass | Apache-2.0 | |
| Graphsignalgraphsignal/graphsignal | 257 | — | ~6.3k | Automated safety check: Pass | Apache-2.0 | |
| One EvalOpenDCAI/One-Eval | 165 | — | ~2.4k | Automated safety check: Pass | Apache-2.0 |
huggingface/skills
Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.
dstackai/dstack
Use with the dstack skill for model-serving work when the image, serving command, resources, backend/fleet choice, or service behavior is not proven.
NVIDIA/OpenShell
Debug inference clients that use an attached provider and its native endpoint, including hosted APIs and host-local Ollama, vLLM, SGLang, TRT-LLM, LM Studio, or NIM.
graphsignal/graphsignal
Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.
OpenDCAI/One-Eval
驱动 One-Eval 对 API 或本地模型做端到端评测,覆盖纯文本、多模态、代码生成、函数调用和 Agent benchmark。当用户想评测模型在一个或多个 benchmark 上的表现、比较分数、补充 metric,或生成图文评测报告时使用本 skill。
BBuf/AI-Infra-Auto-Driven-SKILLS
Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.
sohu-mptc/FlashRec
Deploy and serve GenRec checkpoints with FlashRec (install, serve.sh, SID trie, wide-beam knobs, health/curl, FP8, profiling).
sohu-mptc/FlashRec
Build SID trie catalogs for FlashRec from OpenOneRec RecIF packed mappings.
sohu-mptc/FlashRec
Capture torch.profiler traces on a running FlashRec server. An agent skill from sohu-mptc/FlashRec.
sohu-mptc/FlashRec
给 FlashRec 引擎接入一个新模型架构(新的 HF checkpoint / 非 Qwen3 结构)。涵盖模型定义、权重合并加载、FP8 双路径、融合 kernel 接线、CUDA graph 兼容、精度校验、以及压测+trace 验证闭环。当用户要"增加/支持/接入新模型"时使用。
Categories
Run RecIF beam×concurrency eval and FlashRec vs SGLang/vLLM/TRT-LLM baselines. Recif Eval is an agent skill from sohu-mptc/FlashRec. Run RecIF beam×concurrency eval and FlashRec vs SGLang/vLLM/TRT-LLM baselines.
Recif Eval fits situations like: the user asks to 评测; runsglangflashrecmatrix; benchsglangcompare.
Run `npx skills add sohu-mptc/FlashRec --skill recif-eval -a claude-code`. Or copy the skill folder (.claude/skills/recif-eval in sohu-mptc/FlashRec) into .claude/skills/recif-eval in your project. Claude Code loads it when a task matches its description.
Run `npx skills add sohu-mptc/FlashRec --skill recif-eval -a codex`. Or copy the skill folder (.claude/skills/recif-eval in sohu-mptc/FlashRec) into .agents/skills/recif-eval in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add sohu-mptc/FlashRec --skill recif-eval -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/recif-eval, .gemini/skills/recif-eval, .github/skills/recif-eval and .opencode/skills/recif-eval in your project.
Going by SKILL.md and its folder, Recif Eval needs the command-line tools its instructions call (python and bash). Our summary lists: Python 3; Docker.
SKILL.md names 1 domain. As links in the text: github.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Recif Eval is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 915 tokens (SKILL.md is roughly 3.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Recif Eval: SageMaker Serving Image Selection (huggingface/skills, 11k stars), Dstack Prototyping (dstackai/dstack, 2.3k stars), Debug Inference (NVIDIA/OpenShell, 16k stars) and Graphsignal (graphsignal/graphsignal, 257 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
sohu-mptc (a GitHub organization) maintains it in sohu-mptc/FlashRec, which has 107 GitHub stars. The repository holds 5 skills in this directory. The repository was last updated on October 10, 2026.
Source: sohu-mptc/FlashRec on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.