Agent skill

Recif Eval

by sohu-mptc in sohu-mptc/FlashRec

Run RecIF beam×concurrency eval and FlashRec vs SGLang/vLLM/TRT-LLM baselines.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Recif Eval

skills CLI
$ npx skills add sohu-mptc/FlashRec --skill recif-eval -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install sohu-mptc/FlashRec recif-eval --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/sohu-mptc/FlashRec.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/recif-eval .claude/skills/recif-eval && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
recif-eval
GitHub stars
107
Token cost
~915 tokens
SKILL.md length
170 words
Files
1
Skills in repo
5
Repo updated
First seen
Licence
Apache-2.0

At a glance

Run RecIF beam×concurrency eval and FlashRec vs SGLang/vLLM/TRT-LLM baselines.

  • The user asks to 评测
  • SKILL.md covers 矩阵:FlashRec vs SGLang-master, 单格客户端, 其他引擎 and 精度 / 对齐测试(非检索质量)
  • Calls python and bash
  • Runsglangflashrecmatrix

What it does

Recif Eval is an agent skill from sohu-mptc/FlashRec. Run RecIF beam×concurrency eval and FlashRec vs SGLang/vLLM/TRT-LLM baselines. Use when the user asks to 评测, 对照, 压测, matrix, RecIF, recall@32, invalidrate, runsglangflashrecmatrix, benchsglangcompare, or evalbeammatrix.

Its SKILL.md is about 920 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering LLM inference and serving. It works with SGLang and vLLM. The repository describes itself as: FlashRec is a CUDA-graph engine for generative recommendation: wide beam search (3–5 SID steps, n=50–512+) over a trie-constrained catalog, in-process FP8 serving, and ranked… The licence is Apache-2.0.

When your agent uses it

  • The user asks to 评测
  • Runsglangflashrecmatrix
  • Benchsglangcompare

Example prompts

  • “/recif-eval”

Requirements

  • Python 3
  • Docker

What it can do on your machine

Read from SKILL.md and the folder at commit 1089682. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python
    • bash

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Recif Eval loads about 915 tokens when it runs. Until then it costs about 60 tokens; SKILL.md has 170 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~60
When it runs · the whole SKILL.md, loaded when a task matches
~915

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from sohu-mptc/FlashRec at commit 1089682, republished under its Apache-2.0 licence (© sohu-mptc). 170 words, ~915 tokens.

Download SKILL.mdSave it as .claude/skills/recif-eval/SKILL.md (or your agent's skills folder).
name
recif-eval
description
Run RecIF beam×concurrency eval and FlashRec vs SGLang/vLLM/TRT-LLM baselines. Use when the user asks to 评测, 对照, 压测, matrix, RecIF, recall@32, invalid_rate, run_sglang_flashrec_matrix, bench_sglang_compare, or eval_beam_matrix.

RecIF 评测与 baseline

公开数字用 OneRec-1.7B(快手 OneRec-1.7B × OpenOneRec RecIF-Bench video)。 对照协议与表格见 docs/baselines.zh-CN.md。

不要把不同 workload 的 QPS 直接相除。 OneRec-1.7B prompt 约 2.5k token(~500 条历史 SID); 同一引擎 n=50 并发 8 时约 26 QPS,饱和约 28 QPS。生产短 prompt(约 300 token) n=50 饱和约 220 QPS(1.7B)/ 303 QPS(0.6B);n=1000 为 24.6 QPS(1.7B)/ 32 QPS(0.6B)。见 docs/baselines.zh-CN.md。

SGLang-master 与 SGLang 0801 分列。 SGLang-master 是 PR #31626;0801 是 cswuyg/sglang feature/beam_search_update_0801。写结论、表格、加速比时不要把 两家 QPS 合成一行。

不公平对照警告: FlashRec 默认 SID trie;SGLang-master / SGLang 0801 / vLLM / HF / TensorRT-LLM 在本仓库文档中是开放词表 beam。比吞吐时对齐 beam_width、 max_tokens、模型、硬件;比质量时必须报 invalid_rate。不要引用 FlashRec conc=1 的 recall@32(unique-beam 塌缩)。

矩阵:FlashRec vs SGLang-master

需要两张卡(默认 FLASHREC_GPU=0、SGL_GPU=1),且 $PYTHON 能 import 带 PR #31626 的 SGLang-master。

bash
# 1. catalog(一次)
DATA_DIR=/path/to/OpenOneRec-RecIF/benchmark_data \
  bash scripts/build_catalog.sh

# 2. 速度对照(推荐):beam {50,128,512} × 并发 {1,8,16,32}
# SGLang-master 默认 Docker 镜像 nightly-dev-cu13-20260827-20621aa1,请求走 /generate + beam_width
MODEL_PATH=/path/to/OneRec-1.7B \
DATA_DIR=/path/to/OpenOneRec-RecIF/benchmark_data \
SMOKE=1 bash scripts/bench_sglang_compare.sh

MODEL_PATH=/path/to/OneRec-1.7B \
DATA_DIR=/path/to/OpenOneRec-RecIF/benchmark_data \
bash scripts/bench_sglang_compare.sh

产物:results/onerec_beam_conc_bench_<stamp>/(MATRIX_REPORT.md、 SPEED_COMPARE.md、matrix_summary.json、每格 summary.json / candidates.csv / per_sample_metrics.csv / latencies.jsonl)。 results/ 已 gitignore。从一次 run 重生成对照表:

bash
python scripts/summarize_sglang_compare.py results/onerec_beam_conc_bench_<stamp>

可覆盖:BEAMS、CONCS、SAMPLE_SIZE、FLASHREC_GPU、SGL_GPU、FLASHREC_PORT、 SGL_PORT、PYTHON、SGLANG_LAUNCH、DOCKER、OUTDIR。FlashRec 按 n 自动设 graph / 槽位(50→800、128→2048、≥512→4096),换宽度会重启两侧服务。

FlashRec 侧只需要 SID_VOCAB_FILE(脚本默认 data/catalogs/sid2pid_beamrec_l4.json); 布局从 tokenizer 推断,不必设 SID。SGLang-master 是开放词表,客户端必须打 POST /generate + sampling_params.beam_width(eval_beam_matrix.py --engine sglang)。

单格客户端

服务已在跑时:

bash
python scripts/eval_beam_matrix.py \
  --engine flashrec \
  --server-url http://127.0.0.1:8000 \
  --data-dir /path/to/benchmark_data \
  --catalog data/catalogs/sid2pid_beamrec_l4.json \
  --out-dir /tmp/cell --task video --n 50 --concurrency 8 --sample-size 200

--engine sglang 走 POST /generate(还要 --model-path)。其它引擎走 /v1/chat/completions。看 summary.json 的 qps、invalid_rate、 metrics.recall@32 / ndcg@32。

其他引擎

启动命令在 docs/baselines.zh-CN.md 文末,镜像不随仓库提供。要点:

  • SGLang-master:请求 sampling_params.beam_width(POST /generate)、 --disable-overlap-schedule、--max-running-requests ≥ (n+1) × 并发 (expand 会复制 request-pool 槽位)。速度对照用 scripts/bench_sglang_compare.sh
  • vLLM:use_beam_search=true;--max-logprobs ≥ 2n(n=32 至少 100)
  • TensorRT-LLM:--max_beam_width 与请求 best_of 一致;32 GB 上 n=512 通常不支持

精度 / 对齐测试(非检索质量)

bash
PYTHONPATH=python python -m unittest discover -s tests -v

SGLANG_BEAM_URL=http://127.0.0.1:PORT FLASHREC_URL=http://127.0.0.1:8000 \
  PYTHONPATH=python python -m unittest tests.test_parity -v

FLASHREC_DIFF_MODEL=/path/to/model \
  PYTHONPATH=python python -m unittest tests.test_beam_search_diff -v

© sohu-mptc, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/recif-eval of sohu-mptc/FlashRec.

Open the folder on GitHubat commit 1089682

Compare with similar skills

Recif Eval next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Recif Eval compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Recif Eval this skillsohu-mptc/FlashRec107—~915Automated safety check: PassApache-2.0
SageMaker Serving Image Selectionhuggingface/skills11k1 repos~4.6kAutomated safety check: PassApache-2.0
Dstack Prototypingdstackai/dstack2.3k—~1.6kAutomated safety check: PassMPL-2.0
Debug InferenceNVIDIA/OpenShell16k—~1.9kAutomated safety check: PassApache-2.0
Graphsignalgraphsignal/graphsignal257—~6.3kAutomated safety check: PassApache-2.0
One EvalOpenDCAI/One-Eval165—~2.4kAutomated safety check: PassApache-2.0

Similar skills

  • Official

    Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.

    11k GitHub starsUsed in 1 repo~4.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Dstack Prototyping

    dstackai/dstack

    Use with the dstack skill for model-serving work when the image, serving command, resources, backend/fleet choice, or service behavior is not proven.

    2.3k GitHub stars~1.6k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Debug Inference

    NVIDIA/OpenShell

    Official

    Debug inference clients that use an attached provider and its native endpoint, including hosted APIs and host-local Ollama, vLLM, SGLang, TRT-LLM, LM Studio, or NIM.

    16k GitHub stars~1.9k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Graphsignal

    graphsignal/graphsignal

    Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

    257 GitHub stars~6.3k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • One Eval

    OpenDCAI/One-Eval

    驱动 One-Eval 对 API 或本地模型做端到端评测,覆盖纯文本、多模态、代码生成、函数调用和 Agent benchmark。当用户想评测模型在一个或多个 benchmark 上的表现、比较分数、补充 metric,或生成图文评测报告时使用本 skill。

    165 GitHub stars~2.4k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • LLM Torch Profiler Trace Analysis

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.

    938 GitHub stars~2.8k tokensUpdated 5 days ago
    AI & LLM EngineeringAuto-check passed

More from sohu-mptc/FlashRec

  • Model Deploy

    sohu-mptc/FlashRec

    Deploy and serve GenRec checkpoints with FlashRec (install, serve.sh, SID trie, wide-beam knobs, health/curl, FP8, profiling).

    107 GitHub stars~974 tokensUpdated today
    Auto-check passed
  • Sid Catalog

    sohu-mptc/FlashRec

    Build SID trie catalogs for FlashRec from OpenOneRec RecIF packed mappings.

    107 GitHub stars~468 tokensUpdated today
    Auto-check passed
  • Profile Serving

    sohu-mptc/FlashRec

    Capture torch.profiler traces on a running FlashRec server. An agent skill from sohu-mptc/FlashRec.

    107 GitHub stars~312 tokensUpdated today
    Auto-check passed
  • Add Model

    sohu-mptc/FlashRec

    给 FlashRec 引擎接入一个新模型架构(新的 HF checkpoint / 非 Qwen3 结构)。涵盖模型定义、权重合并加载、FP8 双路径、融合 kernel 接线、CUDA graph 兼容、精度校验、以及压测+trace 验证闭环。当用户要"增加/支持/接入新模型"时使用。

    107 GitHub stars~1.7k tokensUpdated today
    Auto-check passed

Works with

Questions about Recif Eval

What does Recif Eval do?

Run RecIF beam×concurrency eval and FlashRec vs SGLang/vLLM/TRT-LLM baselines. Recif Eval is an agent skill from sohu-mptc/FlashRec. Run RecIF beam×concurrency eval and FlashRec vs SGLang/vLLM/TRT-LLM baselines.

When should I use Recif Eval?

Recif Eval fits situations like: the user asks to 评测; runsglangflashrecmatrix; benchsglangcompare.

How do I install Recif Eval in Claude Code?

Run `npx skills add sohu-mptc/FlashRec --skill recif-eval -a claude-code`. Or copy the skill folder (.claude/skills/recif-eval in sohu-mptc/FlashRec) into .claude/skills/recif-eval in your project. Claude Code loads it when a task matches its description.

How do I install Recif Eval in Codex?

Run `npx skills add sohu-mptc/FlashRec --skill recif-eval -a codex`. Or copy the skill folder (.claude/skills/recif-eval in sohu-mptc/FlashRec) into .agents/skills/recif-eval in your project. Codex loads it when a task matches its description.

Can I use Recif Eval in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add sohu-mptc/FlashRec --skill recif-eval -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/recif-eval, .gemini/skills/recif-eval, .github/skills/recif-eval and .opencode/skills/recif-eval in your project.

What does Recif Eval need to run?

Going by SKILL.md and its folder, Recif Eval needs the command-line tools its instructions call (python and bash). Our summary lists: Python 3; Docker.

Does Recif Eval access the network?

SKILL.md names 1 domain. As links in the text: github.com. This is read from the text; nothing was executed.

Is Recif Eval safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Recif Eval use?

Recif Eval is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Recif Eval use?

About 915 tokens (SKILL.md is roughly 3.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Recif Eval?

Skills that share tags, products or a category with Recif Eval: SageMaker Serving Image Selection (huggingface/skills, 11k stars), Dstack Prototyping (dstackai/dstack, 2.3k stars), Debug Inference (NVIDIA/OpenShell, 16k stars) and Graphsignal (graphsignal/graphsignal, 257 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Recif Eval?

sohu-mptc (a GitHub organization) maintains it in sohu-mptc/FlashRec, which has 107 GitHub stars. The repository holds 5 skills in this directory. The repository was last updated on October 10, 2026.

Source: sohu-mptc/FlashRec on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.