LLM Benchmarking with lm-evaluation-harness
Orchestra-Research/AI-Research-SKILLs
Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.
Sets up and runs side-by-side model evaluations with frozen questions, isolated tools and traceable results, using your own question bank or benchmark subsets.
SKILL.md written in Chinese; this summary is our English description.
$ npx skills add liucongg/liucong-skills --skill liucong-model-eval -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install liucongg/liucong-skills liucong-model-eval --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/liucongg/liucong-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/liucong-model-eval .claude/skills/liucong-model-eval && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "liucong-model-eval" agent skill from https://github.com/liucongg/liucong-skills/tree/main/skills/liucong-model-eval into .claude/skills/liucong-model-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "liucong-model-eval", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/liucongg/liucong-skills/tree/main/skills/liucong-model-evalType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add liucongg/liucong-skills --skill liucong-model-eval -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install liucongg/liucong-skills liucong-model-eval --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/liucongg/liucong-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/liucong-model-eval .agents/skills/liucong-model-eval && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "liucong-model-eval" agent skill from https://github.com/liucongg/liucong-skills/tree/main/skills/liucong-model-eval into .agents/skills/liucong-model-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "liucong-model-eval", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add liucongg/liucong-skills --skill liucong-model-eval -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install liucongg/liucong-skills liucong-model-eval --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/liucongg/liucong-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/liucong-model-eval .cursor/skills/liucong-model-eval && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "liucong-model-eval" agent skill from https://github.com/liucongg/liucong-skills/tree/main/skills/liucong-model-eval into .cursor/skills/liucong-model-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "liucong-model-eval", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/liucongg/liucong-skills.git --path skills/liucong-model-eval--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add liucongg/liucong-skills --skill liucong-model-eval -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install liucongg/liucong-skills liucong-model-eval --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/liucongg/liucong-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/liucong-model-eval .gemini/skills/liucong-model-eval && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "liucong-model-eval" agent skill from https://github.com/liucongg/liucong-skills/tree/main/skills/liucong-model-eval into .gemini/skills/liucong-model-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "liucong-model-eval", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install liucongg/liucong-skills liucong-model-evalInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add liucongg/liucong-skills --skill liucong-model-eval -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/liucongg/liucong-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/liucong-model-eval .github/skills/liucong-model-eval && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "liucong-model-eval" agent skill from https://github.com/liucongg/liucong-skills/tree/main/skills/liucong-model-eval into .github/skills/liucong-model-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "liucong-model-eval", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add liucongg/liucong-skills --skill liucong-model-eval -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install liucongg/liucong-skills liucong-model-eval --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/liucongg/liucong-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/liucong-model-eval .opencode/skills/liucong-model-eval && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "liucong-model-eval" agent skill from https://github.com/liucongg/liucong-skills/tree/main/skills/liucong-model-eval into .opencode/skills/liucong-model-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "liucong-model-eval", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
liucong-model-evalSets up and runs side-by-side model evaluations with frozen questions, isolated tools and traceable results, using your own question bank or benchmark subsets.
This skill sets up and runs traceable side-by-side model evaluations with fixed questions. It separates what a model answered, the generation state and the actual acceptance result, and the orchestrating agent never answers for the model under test or counts a page it fixed as the model's first-round result. It supports subsets of published visual benchmarks, your own question bank and real front-end and back-end tasks. The built-in questions are original demos, not a private question bank or an official leaderboard. The skill text is in Chinese.
Setup runs node scripts/setup.mjs doctor, then connect.mjs, where you type your own key with hidden input so it never lands in config, reports or chat, then an isolation check and per-model connection, vision and tool checks that do not count as formal results. Runs freeze the questions, images, seed, tools and budget so every model sees identical conditions, execute in a sandboxed process with fresh working directories and no network by default, and keep failed attempts. Reports trace each conclusion to a run ID, content hashes, the raw answer and actual actions. The bundled runner supports macOS only.
5 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit d08416a. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/, which the agent can run.
Shell commands in SKILL.md call:
nodeFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Traceable Model Evaluation loads about 752 tokens when it runs, and up to ~6.9k if it reads all its reference files. Until then it costs about 31 tokens; SKILL.md has 123 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from liucongg/liucong-skills at commit d08416a, republished under its Apache-2.0 licence (© liucongg). 123 words, ~752 tokens.
.claude/skills/liucong-model-eval/SKILL.md (or your agent's skills folder). This skill also uses 22 other files; get the full folder from GitHub.用固定题目测试真实接入的模型,区分模型回答、生成状态和实际验收。调度 Agent 不替被测模型答题,也不把自己修好的页面记成模型首轮成绩。本包按刘聪式实测方法组织,内置题是原创演示题,不是刘聪完整私有题库或官方榜单。
所有下列命令从本 Skill 文件夹执行;路径有空格时用双引号引用,或使用程序参数数组。脚本路径相对 Skill,自定义题库路径相对该题库。
node scripts/setup.mjs doctor。未初始化则按初始化文档补软件、运行 init;已有配置保留,不清空用户的 Claude Code 全局配置。node scripts/connect.mjs,由本人隐藏输入自己的 Agent Plan Key。Key 只在连接进程内存里;配置、Skill、题库、报告和截图均不得含真实 Key。不能安全输入时给用户这一步,不让其把 Key 发到聊天。node scripts/runner.mjs isolation-check。越界读写、外网必须拒绝,目录内写、Node、Claude CLI 启动必须通过;失败则停在具体错误,不换成无隔离方式。常用流程:
node scripts/runner.mjs run --cases=connection,visioncheck,toolscheck --model=glm-5.3-flash
node scripts/runner.mjs run --dry-run --tier=simple --count=3
node scripts/runner.mjs run --tier=simple --count=3
node scripts/runner.mjs run --cases=orbit_audio --model=glm-5.3-flash --seconds=1200
node scripts/runner.mjs status
node scripts/runner.mjs export自有/权威题库加 --bank="/题库目录/bank.json";如果配置了非内置题库,准备检查显式传本 Skill 的 assets/demo/bank.json。参数细节和个人偏好见题库文档。
读 方法与验收。每条结论能追溯到 runId、题面/图片哈希、原始回答和实际操作。
© liucongg, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 22 other files (scripts, references, assets) in skills/liucong-model-eval of liucongg/liucong-skills.
Open the folder on GitHubat commit d08416a
Traceable Model Evaluation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Traceable Model Evaluation this skillliucongg/liucong-skills | 248 | — | ~752 | Automated safety check: Pass | Apache-2.0 | |
| LLM Benchmarking with lm-evaluation-harnessOrchestra-Research/AI-Research-SKILLs | 13k | 8 repos | ~3k | Automated safety check: Pass | MIT | |
| Caveman Experiment ManagerJuliusBrussee/caveman | 111k | 1 repos | ~975 | Automated safety check: Pass | Apache-2.0 | |
| Caveman Optimization EvaluatorJuliusBrussee/caveman | 111k | 1 repos | ~1.2k | Automated safety check: Pass | Apache-2.0 | |
| Evals Contextzgsm-ai/costrict | 4.5k | 1 repos | ~1.9k | Automated safety check: Pass | Apache-2.0 | |
| Hugging Face Local Model Evalshuggingface/skills | 11k | 2 repos | ~1.6k | Automated safety check: Pass | Apache-2.0 |
Orchestra-Research/AI-Research-SKILLs
Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.
JuliusBrussee/caveman
Reads the state and results of Caveman Cloud experiments and reports one recommendation or a block, without changing an experiment's lifecycle itself.
JuliusBrussee/caveman
Turns a Caveman report-only optimization observation into one minimal code change and a paired baseline evaluation, after the operator picks which to pursue.
zgsm-ai/costrict
Provides context about the CoStrict evals system structure in this monorepo.
huggingface/skills
Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.
chatboxai/chatbox
Runs and debugs evaluations of how Chatbox models answer questions about large attached files, using synthetic and real long-document fixtures.
liucongg/liucong-skills
Turns a short brief into a researched, distinctive website: lock the intent with at most one question, research live reference sites, source materials, then build and verify.
liucongg/liucong-skills
Expands one sentence into a prompt package for painterly 3D-to-2D short films: story, 15-second segments, characters, scenes, Midjourney storyboards and Seedance video prompts.
liucongg/liucong-skills
Maintains an LLM Wiki in a Feishu knowledge base: initial setup, ingesting sources and articles, answering queries, health checks and entry upkeep.
liucongg/liucong-skills
Generates, rewrites, critiques and ranks titles for Chinese WeChat official account articles, grounded in the article's real content and any history data supplied.
Categories
Sets up and runs side-by-side model evaluations with frozen questions, isolated tools and traceable results, using your own question bank or benchmark subsets. This skill sets up and runs traceable side-by-side model evaluations with fixed questions. It separates what a model answered, the generation state and the actual acceptance result, and the orchestrating agent never answers for the model under test or counts a page it fixed as the model's first-round result.
Traceable Model Evaluation fits situations like: setting up an environment to compare models on the same questions; running your own question bank against several models; testing models on a subset of a published visual benchmark; assembling an evaluation report that traces each conclusion to a run.
Run `npx skills add liucongg/liucong-skills --skill liucong-model-eval -a claude-code`. Or copy the skill folder (skills/liucong-model-eval in liucongg/liucong-skills) into .claude/skills/liucong-model-eval in your project. Claude Code loads it when a task matches its description.
Run `npx skills add liucongg/liucong-skills --skill liucong-model-eval -a codex`. Or copy the skill folder (skills/liucong-model-eval in liucongg/liucong-skills) into .agents/skills/liucong-model-eval in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add liucongg/liucong-skills --skill liucong-model-eval -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/liucong-model-eval, .gemini/skills/liucong-model-eval, .github/skills/liucong-model-eval and .opencode/skills/liucong-model-eval in your project.
Going by SKILL.md and its folder, Traceable Model Evaluation needs the command-line tools its instructions call (node). Our summary lists: Node to run the setup, connect and runner scripts; macOS for the bundled isolated runner; A model API key you enter yourself at a terminal prompt.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Traceable Model Evaluation is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 752 tokens (SKILL.md is roughly 3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 6.2k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Traceable Model Evaluation: LLM Benchmarking with lm-evaluation-harness (Orchestra-Research/AI-Research-SKILLs, 13k stars), Caveman Experiment Manager (JuliusBrussee/caveman, 111k stars), Caveman Optimization Evaluator (JuliusBrussee/caveman, 111k stars) and Evals Context (zgsm-ai/costrict, 4.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
liucongg (a GitHub user) maintains it in liucongg/liucong-skills, which has 248 GitHub stars. The repository holds 5 skills in this directory. The repository was last updated on September 8, 2026.
Source: liucongg/liucong-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.