Agent skill

Traceable Model Evaluation

by liucongg in liucongg/liucong-skills

Sets up and runs side-by-side model evaluations with frozen questions, isolated tools and traceable results, using your own question bank or benchmark subsets.

Apache-2.0Auto-check passedAI & LLM Engineering

SKILL.md written in Chinese; this summary is our English description.

Install Traceable Model Evaluation

skills CLI
$ npx skills add liucongg/liucong-skills --skill liucong-model-eval -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install liucongg/liucong-skills liucong-model-eval --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/liucongg/liucong-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/liucong-model-eval .claude/skills/liucong-model-eval && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
liucong-model-eval
GitHub stars
248
Token cost
~752 tokens
SKILL.md length
123 words
Files
23 (incl. scripts, references, assets)
Skills in repo
5
Repo updated
First seen
Licence
Apache-2.0

At a glance

Sets up and runs side-by-side model evaluations with frozen questions, isolated tools and traceable results, using your own question bank or benchmark subsets.

  • Works in 5 steps: 运行 node scripts/setup.mjs… → 用户在终端执行 node… → 运行 node scripts/runner.mjs… → …
  • Setting up an environment to compare models on the same questions
  • SKILL.md covers 先识别当前情况, 初始化与凭据, 冻结题目,再执行 and 验收与交付
  • Calls node

What it does

This skill sets up and runs traceable side-by-side model evaluations with fixed questions. It separates what a model answered, the generation state and the actual acceptance result, and the orchestrating agent never answers for the model under test or counts a page it fixed as the model's first-round result. It supports subsets of published visual benchmarks, your own question bank and real front-end and back-end tasks. The built-in questions are original demos, not a private question bank or an official leaderboard. The skill text is in Chinese.

Setup runs node scripts/setup.mjs doctor, then connect.mjs, where you type your own key with hidden input so it never lands in config, reports or chat, then an isolation check and per-model connection, vision and tool checks that do not count as formal results. Runs freeze the questions, images, seed, tools and budget so every model sees identical conditions, execute in a sandboxed process with fresh working directories and no network by default, and keep failed attempts. Reports trace each conclusion to a run ID, content hashes, the raw answer and actual actions. The bundled runner supports macOS only.

When your agent uses it

  • Setting up an environment to compare models on the same questions
  • Running your own question bank against several models
  • Testing models on a subset of a published visual benchmark
  • Assembling an evaluation report that traces each conclusion to a run

Example prompts

  • “Set up the model evaluation environment from scratch and run the connection checks.”
  • “Do a dry run of the simple tier and show me which questions and models are selected.”
  • “Compare two models on my question bank in ./my-bank and write up the results.”

Requirements

  • Node to run the setup, connect and runner scripts
  • macOS for the bundled isolated runner
  • A model API key you enter yourself at a terminal prompt

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. 运行 node scripts/setup.mjs doctor。未初始化则按初始化文档补软件、运行 init;已有配置保留,不清空用户的 Claude Code 全局配置。
  2. 用户在终端执行 node scripts/connect.mjs,由本人隐藏输入自己的 Agent Plan Key。Key 只在连接进程内存里;配置、Skill、题库、报告和截图均不得含真实 Key。不能安全输入时给用户这一步,不让其把 Key 发到聊天。
  3. 运行 node scripts/runner.mjs isolation-check。越界读写、外网必须拒绝,目录内写、Node、Claude CLI 启动必须通过;失败则停在具体错误,不换成无隔离方式。
  4. 对每个模型先跑 connection;要用图片再跑 visioncheck,要做代码题再跑 toolscheck。准备检查不计正式测评。工具检查核对真实 tool_result,模型自述不算证据。视觉探针失败不直接断言模型不支持视觉,先检查接口错误与原始回答。
  5. 连接关闭或重启后重新输入 Key,并重新做该连接的准备检查。更换模型名单后重启连接。接口拒绝、模型名错误或额度不足不自动切模型。

What it can do on your machine

Read from SKILL.md and the folder at commit d08416a. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/, which the agent can run.

    Shell commands in SKILL.md call:

    • node

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Traceable Model Evaluation loads about 752 tokens when it runs, and up to ~6.9k if it reads all its reference files. Until then it costs about 31 tokens; SKILL.md has 123 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~31
When it runs · the whole SKILL.md, loaded when a task matches
~752
With references · SKILL.md plus every file in references/, read only if the agent opens them
~6.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from liucongg/liucong-skills at commit d08416a, republished under its Apache-2.0 licence (© liucongg). 123 words, ~752 tokens.

Download SKILL.mdSave it as .claude/skills/liucong-model-eval/SKILL.md (or your agent's skills folder). This skill also uses 22 other files; get the full folder from GitHub.
name
liucong-model-eval
description
初始化并执行可追溯的模型对照测评,支持权威视觉 benchmark 子集、用户自有题库和前后端真实任务;固定题面、隔离工具、保留首轮产物并验收。用于从零搭建测评环境、按个人题库跑测、同题比较模型或整理测评报告。
metadata.version
2.0.0

刘聪模型测评

用固定题目测试真实接入的模型,区分模型回答、生成状态和实际验收。调度 Agent 不替被测模型答题,也不把自己修好的页面记成模型首轮成绩。本包按刘聪式实测方法组织,内置题是原创演示题,不是刘聪完整私有题库或官方榜单。

先识别当前情况

  • 第一次使用 / 缺软件 / 缺 Key:读 初始化,从环境检查开始。不要假设别人的电脑已经配置好;不要沿用演示机路径。随包执行器支持 macOS,其他系统未提供已验证适配器,不降级为无隔离执行。
  • 用户已经有题库或偏好:先用用户指定的题面、图片和评分方式,读 自有题库与偏好。不擅自改题,不用演示题顶替。
  • 用户要找权威题 / benchmark:读 题源目录,从作者/机构官方渠道获取原题,冻结版本、题号、图片和答案。少量抽测称“某 benchmark 子集”,不得标成完整 benchmark 分数。
  • 只说“简单测一下”:沿用已保存偏好;没有偏好时用内置简单演示题,默认最多3题、先1个模型。手机指令不自动扩大成长任务。

所有下列命令从本 Skill 文件夹执行;路径有空格时用双引号引用,或使用程序参数数组。脚本路径相对 Skill,自定义题库路径相对该题库。

初始化与凭据

  1. 运行 node scripts/setup.mjs doctor。未初始化则按初始化文档补软件、运行 init;已有配置保留,不清空用户的 Claude Code 全局配置。
  2. 用户在终端执行 node scripts/connect.mjs,由本人隐藏输入自己的 Agent Plan Key。Key 只在连接进程内存里;配置、Skill、题库、报告和截图均不得含真实 Key。不能安全输入时给用户这一步,不让其把 Key 发到聊天。
  3. 运行 node scripts/runner.mjs isolation-check。越界读写、外网必须拒绝,目录内写、Node、Claude CLI 启动必须通过;失败则停在具体错误,不换成无隔离方式。
  4. 对每个模型先跑 connection;要用图片再跑 visioncheck,要做代码题再跑 toolscheck。准备检查不计正式测评。工具检查核对真实 tool_result,模型自述不算证据。视觉探针失败不直接断言模型不支持视觉,先检查接口错误与原始回答。
  5. 连接关闭或重启后重新输入 Key,并重新做该连接的准备检查。更换模型名单后重启连接。接口拒绝、模型名错误或额度不足不自动切模型。

冻结题目,再执行

  • 先用 bank validate 检查题库;用 run --dry-run 列出本轮题号、模型、预算,不调用模型。
  • 模型对照使用同题面、同图片字节/顺序、同初始工程、同工具、同预算。固定 seed 和题号;保留全部选中题,不看结果后换题。不同题库、工具条件或预算分组报告。
  • 支持一次 run --models=模型A,模型B 顺序执行,避免抢同一本地服务端口。默认不自动重跑已有同条件任务;补跑需明确理由并加 --repeat --reason="原因",保留失败记录。
  • 简单问答禁用全部工具;代码题限 Read、Write、Edit、Bash,默认无外网。每题新工作目录与 Claude 配置,禁用全局 Skill/MCP/记忆/历史发现。不能使用跳过权限或关闭隔离参数。
  • 题库答案、解析、评分脚本不复制进被测模型目录。图片作为真实 image 内容传入,按 Image 1…顺序记录;不能用调度者的图像描述替代视觉输入。
  • 隔离连接给每题单独的短期令牌,绑定单个模型;被测进程不接触真实上游 Key。系统运行库仍可读,本适配器是进程沙箱,不宣称是虚拟机。
  • 生成的命令和网页视为待验收产物;后端在同样的受限环境跑测试,不能直接在宿主机执行不受限代码。

常用流程:

sh
node scripts/runner.mjs run --cases=connection,visioncheck,toolscheck --model=glm-5.3-flash
node scripts/runner.mjs run --dry-run --tier=simple --count=3
node scripts/runner.mjs run --tier=simple --count=3
node scripts/runner.mjs run --cases=orbit_audio --model=glm-5.3-flash --seconds=1200
node scripts/runner.mjs status
node scripts/runner.mjs export

自有/权威题库加 --bank="/题库目录/bank.json";如果配置了非内置题库,准备检查显式传本 Skill 的 assets/demo/bank.json。参数细节和个人偏好见题库文档。

验收与交付

读 方法与验收。每条结论能追溯到 runId、题面/图片哈希、原始回答和实际操作。

  • complete 只表示会话结束,不等于答案正确或页面通过。文件缺失、超时但有产物、接口错误、人工复核、未测分别记录。
  • 选择题仅自动认单个选项;解释很长或答案歧义需复核,不能从任意字母猜答案。开放题按官方评分器或预先约定 rubric;本包的小样本严格匹配不是官方榜单评分器。
  • 前端实际打开并操作关键功能,视觉和功能分开评。手机响应式记录真实 innerWidth、scrollWidth。截图只代表当时画面;不能把按钮状态当听感、把隐藏DOM当可见弹层。
  • 每个重点案例保留1张首屏和1张关键交互/问题截图即可;过程图备档,不要求全部入稿。没浏览器能力就写未做GUI验收。
  • 首轮不覆盖;用户要求修复时才另存修复版,记操作者和哈希。用户说停止/不修就收敛。
  • export 输出 CSV、JSON、原始 Prompt、图片、事件和产物;再组织“结论 → 同题条件 → 案例证据 → 限制/失败”。自有私密题库默认只在本地使用,不随公开 Skill 打包。
  • 飞书/云电脑是可选交付:有权限才写真实文档,没权限先交本地文件;没有真实验证不宣称已自动同步或手机远控成功。文章语气可另用写作 Skill,本 Skill 执行不依赖它。

© liucongg, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 22 other files (scripts, references, assets) in skills/liucong-model-eval of liucongg/liucong-skills.

  • SKILL.md
  • agents/openai.yaml
  • assets/custom-example/bank.json
  • assets/demo/bank.json
  • assets/demo/board-source.html
  • assets/demo/clock.png
  • assets/demo/illusion.png
  • references/benchmark-sources.md
  • references/initialization.md
  • references/local-machine.md
  • references/methodology.md
  • references/question-bank.md
  • references/visual-prompts/board-redesign.md
  • references/visual-prompts/orbit-audio.md
  • scripts
  • … and 8 more

Open the folder on GitHubat commit d08416a

Compare with similar skills

Traceable Model Evaluation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Traceable Model Evaluation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Traceable Model Evaluation this skillliucongg/liucong-skills248—~752Automated safety check: PassApache-2.0
LLM Benchmarking with lm-evaluation-harnessOrchestra-Research/AI-Research-SKILLs13k8 repos~3kAutomated safety check: PassMIT
Caveman Experiment ManagerJuliusBrussee/caveman111k1 repos~975Automated safety check: PassApache-2.0
Caveman Optimization EvaluatorJuliusBrussee/caveman111k1 repos~1.2kAutomated safety check: PassApache-2.0
Evals Contextzgsm-ai/costrict4.5k1 repos~1.9kAutomated safety check: PassApache-2.0
Hugging Face Local Model Evalshuggingface/skills11k2 repos~1.6kAutomated safety check: PassApache-2.0

Similar skills

  • LLM Benchmarking with lm-evaluation-harness

    Orchestra-Research/AI-Research-SKILLs

    Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.

    13k GitHub starsUsed in 8 repos~3k tokens
    AI & LLM EngineeringAuto-check passed
  • Caveman Experiment Manager

    JuliusBrussee/caveman

    Reads the state and results of Caveman Cloud experiments and reports one recommendation or a block, without changing an experiment's lifecycle itself.

    111k GitHub starsUsed in 1 repo~975 tokens
    AI & LLM EngineeringAuto-check passed
  • Caveman Optimization Evaluator

    JuliusBrussee/caveman

    Turns a Caveman report-only optimization observation into one minimal code change and a paired baseline evaluation, after the operator picks which to pursue.

    111k GitHub starsUsed in 1 repo~1.2k tokens
    AI & LLM EngineeringAuto-check passed
  • Evals Context

    zgsm-ai/costrict

    Provides context about the CoStrict evals system structure in this monorepo.

    4.5k GitHub starsUsed in 1 repo~1.9k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.

    11k GitHub starsUsed in 2 repos~1.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Chatbox Session RAG Eval

    chatboxai/chatbox

    Runs and debugs evaluations of how Chatbox models answer questions about large attached files, using synthetic and real long-document fixtures.

    42k GitHub stars~758 tokensUpdated 16 days ago
    AI & LLM EngineeringAuto-check passed

More from liucongg/liucong-skills

  • Design-Led Website Builder

    liucongg/liucong-skills

    Turns a short brief into a researched, distinctive website: lock the intent with at most one question, research live reference sites, source materials, then build and verify.

    248 GitHub stars~1.3k tokensUpdated 1 mo ago
    Auto-check passed
  • Painterly 3D2 Cinema Prompts

    liucongg/liucong-skills

    Expands one sentence into a prompt package for painterly 3D-to-2D short films: story, 15-second segments, characters, scenes, Midjourney storyboards and Seedance video prompts.

    248 GitHub stars~1.8k tokensUpdated 1 mo ago
    Auto-check passed
  • LLM Wiki Operations

    liucongg/liucong-skills

    Maintains an LLM Wiki in a Feishu knowledge base: initial setup, ingesting sources and articles, answering queries, health checks and entry upkeep.

    248 GitHub stars~898 tokensUpdated 1 mo ago
    Auto-check passed
  • WeChat Article Title Strategist

    liucongg/liucong-skills

    Generates, rewrites, critiques and ranks titles for Chinese WeChat official account articles, grounded in the article's real content and any history data supplied.

    248 GitHub stars~471 tokensUpdated 1 mo ago
    Auto-check passed

Questions about Traceable Model Evaluation

What does Traceable Model Evaluation do?

Sets up and runs side-by-side model evaluations with frozen questions, isolated tools and traceable results, using your own question bank or benchmark subsets. This skill sets up and runs traceable side-by-side model evaluations with fixed questions. It separates what a model answered, the generation state and the actual acceptance result, and the orchestrating agent never answers for the model under test or counts a page it fixed as the model's first-round result.

When should I use Traceable Model Evaluation?

Traceable Model Evaluation fits situations like: setting up an environment to compare models on the same questions; running your own question bank against several models; testing models on a subset of a published visual benchmark; assembling an evaluation report that traces each conclusion to a run.

How do I install Traceable Model Evaluation in Claude Code?

Run `npx skills add liucongg/liucong-skills --skill liucong-model-eval -a claude-code`. Or copy the skill folder (skills/liucong-model-eval in liucongg/liucong-skills) into .claude/skills/liucong-model-eval in your project. Claude Code loads it when a task matches its description.

How do I install Traceable Model Evaluation in Codex?

Run `npx skills add liucongg/liucong-skills --skill liucong-model-eval -a codex`. Or copy the skill folder (skills/liucong-model-eval in liucongg/liucong-skills) into .agents/skills/liucong-model-eval in your project. Codex loads it when a task matches its description.

Can I use Traceable Model Evaluation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add liucongg/liucong-skills --skill liucong-model-eval -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/liucong-model-eval, .gemini/skills/liucong-model-eval, .github/skills/liucong-model-eval and .opencode/skills/liucong-model-eval in your project.

What does Traceable Model Evaluation need to run?

Going by SKILL.md and its folder, Traceable Model Evaluation needs the command-line tools its instructions call (node). Our summary lists: Node to run the setup, connect and runner scripts; macOS for the bundled isolated runner; A model API key you enter yourself at a terminal prompt.

Does Traceable Model Evaluation access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Traceable Model Evaluation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Traceable Model Evaluation use?

Traceable Model Evaluation is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Traceable Model Evaluation use?

About 752 tokens (SKILL.md is roughly 3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 6.2k tokens, read only when the agent opens those files.

What are the alternatives to Traceable Model Evaluation?

Skills that share tags, products or a category with Traceable Model Evaluation: LLM Benchmarking with lm-evaluation-harness (Orchestra-Research/AI-Research-SKILLs, 13k stars), Caveman Experiment Manager (JuliusBrussee/caveman, 111k stars), Caveman Optimization Evaluator (JuliusBrussee/caveman, 111k stars) and Evals Context (zgsm-ai/costrict, 4.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Traceable Model Evaluation?

liucongg (a GitHub user) maintains it in liucongg/liucong-skills, which has 248 GitHub stars. The repository holds 5 skills in this directory. The repository was last updated on September 8, 2026.

Source: liucongg/liucong-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.