Agent skill

Run Eval

by ai-twinkle in ai-twinkle/Eval

用 Twinkle Eval 跑一次評測——建立 config.yaml、下載或指定資料集、驗證設定、執行並讀結果。當使用者說「跑 benchmark」「跑評測」「建 config」「evaluate 這個模型」「twinkle-eval 怎麼跑」「評測結果怎麼看」「為什麼分數是 0」時使用。涵蓋 15 種 evaluationmethod 的 config 差異、27 個內建可下載…

MITAuto-check passedAI & LLM Engineering

Install Run Eval

skills CLI
$ npx skills add ai-twinkle/Eval --skill run-eval -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ai-twinkle/Eval run-eval --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ai-twinkle/Eval.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/run-eval .claude/skills/run-eval && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
run-eval
GitHub stars
117
Token cost
~1.7k tokens
SKILL.md length
342 words
Files
2 (incl. references)
Skills in repo
2
Repo updated
First seen
Licence
MIT

At a glance

用 Twinkle Eval 跑一次評測——建立 config.yaml、下載或指定資料集、驗證設定、執行並讀結果。當使用者說「跑 benchmark」「跑評測」「建 config」「evaluate 這個模型」「twinkle-eval 怎麼跑」「評測結果怎麼看」「為什麼分數是 0」時使用。涵蓋 15 種 evaluationmethod 的 config 差異、27 個內建可下載…

  • Works in 5 steps: 產生 config → 準備資料集 → 本機 config 的命名規則(重要) → …
  • AI & LLM Engineering work in your project
  • SKILL.md covers 前提:本專案不啟動模型服務, 1. 產生 config, 2. 準備資料集 and 3. 本機 config 的命名規則(重要), plus 4 more sections
  • Calls jq and git

What it does

Run Eval is an agent skill from ai-twinkle/Eval. 用 Twinkle Eval 跑一次評測——建立 config.yaml、下載或指定資料集、驗證設定、執行並讀結果。當使用者說「跑 benchmark」「跑評測」「建 config」「evaluate 這個模型」「twinkle-eval 怎麼跑」「評測結果怎麼看」「為什麼分數是 0」時使用。涵蓋 15 種 evaluationmethod 的 config 差異、27 個內建可下載 benchmark、--validate / --dry-run / --resume,以及用 unparsedrate 診斷 extractor 失效。

Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/config-reference.md`).

It sits in AI & LLM Engineering. It works with OpenAI. The repository describes itself as: High-performance LLM evaluation framework with parallel API calls — up to 17× faster than sequential tools. Supports box, math, and logit-based evaluation. The licence is MIT.

When your agent uses it

  • AI & LLM Engineering work in your project

Example prompts

  • “/run-eval”

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. 產生 config
  2. 準備資料集
  3. 本機 config 的命名規則(重要)
  4. 執行
  5. 讀結果

What it can do on your machine

Read from SKILL.md and the folder at commit 608273c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • jq
    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Run Eval loads about 1.7k tokens when it runs, and up to ~4k if it reads all its reference files. Until then it costs about 71 tokens; SKILL.md has 342 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~71
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from ai-twinkle/Eval at commit 608273c, republished under its MIT licence (© ai-twinkle). 342 words, ~1,667 tokens.

Download SKILL.mdSave it as .claude/skills/run-eval/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
run-eval
description
用 Twinkle Eval 跑一次評測——建立 config.yaml、下載或指定資料集、驗證設定、執行並讀結果。當使用者說「跑 benchmark」「跑評測」「建 config」「evaluate 這個模型」「twinkle-eval 怎麼跑」「評測結果怎麼看」「為什麼分數是 0」時使用。涵蓋 15 種 evaluation_method 的 config 差異、27 個內建可下載 benchmark、--validate / --dry-run / --resume,以及用 unparsed_rate 診斷 extractor 失效。

跑一次評測

前提:本專案不啟動模型服務

Twinkle Eval 只做一件事——拿題目去呼叫已經在外部運行的 OpenAI 相容端點。 模型要自己先起好(vLLM、Ollama、OpenAI、NVIDIA Build 都行),再把 base_url 填進 config。

端點沒回應時,本專案的正確行為是依 max_retries 重試後報錯退出,不會也不該嘗試 重啟服務(CLAUDE.md 原則 G)。

1. 產生 config

bash
twinkle-eval --init                  # 列出全部 11 個範本
twinkle-eval --init multiple_choice  # 產生 configs/multiple_choice.yaml
twinkle-eval --init all              # 全部產生到 configs/

範本存放在 twinkle_eval/templates/,--init 直接掃該目錄。

config 骨架
yaml
llm_api:
  base_url: "http://localhost:8000/v1"   # 必填
  api_key: "EMPTY"                       # 必填(本地 vLLM 隨便填)
  api_rate_limit: -1                     # QPS,-1 為不限
  max_retries: 3
  timeout: 600
  disable_ssl_verify: false

model:
  name: "my-model"                       # 必填,會寫進結果路徑與紀錄
  temperature: 0.0
  top_p: 0.9
  max_tokens: 4096
  extra_body:                            # 傳給 API 的額外參數

evaluation:
  dataset_paths:                         # 必填,list(即使只有一個)
    - "datasets/example/tmmluplus/"
  evaluation_method: "box"               # 必填,見下表
  repeat_runs: 1                         # >1 時算平均與標準差
  shuffle_options: false                 # 選項隨機排列

logging:
  level: "INFO"

每個 evaluation_method 的必填欄位與 strategy_config 參數不同, 完整對照見 references/config-reference.md。

挑 evaluation_method
方法用在備註
pattern選擇題,通用首選正則比對,含中英文預設模式
box選擇題,推理模型提取 \boxed{} / \box{};需要 system_prompt
logit多選項題比較各選項 log-probability,不依賴輸出格式;選項數不限
math數學推理\boxed{} + MathRuler;需 [math]
regex_matchBBH 之類自由格式⚠️ system_prompt 不生效,見下方說明
custom_regex自訂格式必須設 strategy_config.patterns
ifeval / ifbench指令遵循需 [ifeval] / [ifbench] + nltk 資料
bfcl_fc / bfcl_prompt函式呼叫FC 走 tools API,prompt 走注入
niah長文本大海撈針
ragasRAG 品質
text2sqlText-to-SQL需設 text2sql_db_base_path
asr語音辨識需 [asr];llm_api.type: whisper 或多模態
vision_mcq視覺多選題需 [vision](縮放用)

⚠️ evaluation.system_prompt 只對 box 與 math 生效。 models/openai.py 的 _build_messages() 以白名單決定是否送出 system message (method in {"box", "math"}),其他方法即使在 config 設了 system_prompt 也不會進 request。 regex_match 這類需要指定輸出格式的方法,格式要求必須寫進資料集的 question 欄位裡。 (專案的 templates/regex_match.yaml 目前有同樣的誤導,見 #144)

bash
twinkle-eval --list-strategies   # 執行期確認可用方法

2. 準備資料集

bash
twinkle-eval --download-dataset list          # 列出 27 個內建 benchmark
twinkle-eval --download-dataset mmlu          # 短名稱
twinkle-eval --download-dataset tmmluplus gsm8k
twinkle-eval --download-dataset all

內建短名稱:mmlu mmlu_pro mmlu_redux tmmluplus supergpqa gpqa formosa_bench gsm8k aime2025 bbh ifeval ifbench bfcl needlebench longbench wikieval librispeech aishell1 fleurs common_voice mmbench mmstar mmmu pope spider bird spider2_lite

gpqa 是 gated dataset,會互動式要求 HuggingFace token。

先用 example 資料集驗證流程跑得通再下載完整 benchmark: datasets/example/ 底下每個 benchmark 都有 10–30 筆的子集,跑一次只要幾秒。

3. 本機 config 的命名規則(重要)

含真實 API 金鑰的 config 絕對不得 commit(CLAUDE.md 原則 E)。這三種前綴/後綴已被 .gitignore 全局排除,本機測試一律用其中之一:

config_local_*.yaml
config_test_*.yaml
*.local.yaml

寫入任何含金鑰的 config 前,先確認該路徑已在 .gitignore 中。不確定就檢查:

bash
git check-ignore -v config_local_myrun.yaml   # 有輸出 = 已被忽略
git diff --staged | grep -i "api_key"         # commit 前確認

4. 執行

bash
twinkle-eval --validate --config config_local_myrun.yaml   # 只驗設定與資料集,不呼叫 API
twinkle-eval --dry-run  --config config_local_myrun.yaml   # 顯示評測計畫,不呼叫 API
twinkle-eval --config config_local_myrun.yaml              # 正式跑
twinkle-eval --config config_local_myrun.yaml --export json csv html excel

永遠先跑 --validate 再跑正式評測。 資料集路徑錯、缺必填欄位、格式不符都會在這一步 抓到,省下一整輪 API 費用與時間。

中斷後續跑(⚠️ 目前失效,見 #145):

bash
twinkle-eval --resume 20260825_1430 --config config_local_myrun.yaml

5. 讀結果

results/
├── results_{timestamp}.json                 # 整體摘要
└── eval_results_{timestamp}_run{N}.jsonl    # 各題明細(append 模式)

摘要含 dataset_results(各資料集的 average_accuracy、average_std、 average_pass_at_k、total_unparsed_count)與 duration_seconds。 config 欄位是移除 api_key 後的設定。

明細每行含 question_id / sample_id / question / correct_answer / predicted_answer / is_correct / llm_output / llm_reasoning_output / token 用量。 各路徑另有增補欄位(logit 的 logprob_scores、ifeval 的四個指標、vision_mcq 的 image_path、asr 的 wer / cer)。

⚠️ 實際輸出沒有 file 與 timestamp 欄位(CLAUDE.md §10 的規範與實作不符)。 這也讓 --resume 目前無法運作:它以 {file}|{question_id} 作為已完成紀錄的 key, file 缺失使得 key 變成 |{idx},與比對端的 {檔案路徑}| 永遠不匹配, 結果是一題都不會跳過、整輪重跑並 append 出重複列。見 #145。

bash
# 快速看正確率
jq -r '.dataset_results | to_entries[] | "\(.key): \(.value.average_accuracy)"' \
  results/results_*.json

# 撈出所有答錯的題目
jq -c 'select(.is_correct == false) | {question_id, predicted_answer, correct_answer}' \
  results/eval_results_*_run0.jsonl | head

診斷:分數異常低

先看 unparsed_rate——這是判斷「模型答錯」還是「extractor 沒抓到」的關鍵。

bash
# 資料集層級(摘要 JSON 用 average_unparsed_rate,per-file 才叫 unparsed_rate)
jq -r '.dataset_results | to_entries[] | "\(.key): \(.value.average_unparsed_rate)"' \
  results/results_*.json
症狀多半是
unparsed_rate 高(>10%)extractor 沒對上輸出格式
unparsed_rate ≈ 0 但分數低模型是真的答錯
全部 0 分且無 unparsedground truth 欄位或正規化對不上
「所有資料集評測均失敗」資料集路徑、格式,或 API 端點問題

extractor 沒抓到時,撈幾筆 llm_output 出來看實際格式:

bash
jq -r 'select(.predicted_answer == null) | .llm_output' \
  results/eval_results_*_run0.jsonl | head -3

常見成因:

  • box 方法但沒設 system_prompt → 模型不知道要用 \boxed{},自然抓不到
  • 推理模型的 think tag → evaluator 會剝離完整的 <think>...</think> / <reason> / <reasoning> 標籤對;只有結尾 tag 而無開頭 tag 會被視為格式不合格而原樣保留
  • content 為 null(vLLM skip_special_tokens=true)→ 會回退讀 reasoning, 再回退 reasoning_content(vLLM 0.18+ 改名,兩者都支援)
  • 選項超過 4 個(MMLU-Pro A–J、SuperGPQA)→ 確認 extractor 支援多字母選項

效能

並行度由 ThreadPoolExecutor 預設值決定,用 llm_api.api_rate_limit 節流:

  • 本地 vLLM:-1(不限)
  • 有 QPS 限制的商用 API:設成實際上限,否則會大量 429

repeat_runs > 1 會線性放大時間,但能得到標準差——量化模型穩定性時才開。

© ai-twinkle, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (references) in .claude/skills/run-eval of ai-twinkle/Eval.

  • SKILL.md
  • references/config-reference.md

Open the folder on GitHubat commit 608273c

Compare with similar skills

Run Eval next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Run Eval compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Run Eval this skillai-twinkle/Eval117—~1.7kAutomated safety check: PassMIT
Chroma Vector DatabaseOrchestra-Research/AI-Research-SKILLs13k8 repos~2.3kAutomated safety check: PassMIT
CLIP Image-Text MatchingOrchestra-Research/AI-Research-SKILLs13k8 repos~1.7kAutomated safety check: PassMIT
Codebase Managementgiancarloerra/SocratiCode3.3k1 repos~1.8kAutomated safety check: PassAGPL-3.0
Azure AI Projects Python SDKmicrosoft/skills3.1k6 repos~2.8kAutomated safety check: PassMIT
Fine-Tuning ExpertJeffallan/claude-skills12k1 repos~1.7kAutomated safety check: PassMIT

Similar skills

  • Chroma Vector Database

    Orchestra-Research/AI-Research-SKILLs

    Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.

    13k GitHub starsUsed in 8 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed
  • CLIP Image-Text Matching

    Orchestra-Research/AI-Research-SKILLs

    Explains OpenAI's CLIP model for zero-shot image classification, image-text similarity, semantic image search and content moderation, with install steps and code patterns.

    13k GitHub starsUsed in 8 repos~1.7k tokens
    AI & LLM EngineeringAuto-check passed
  • Codebase Management

    giancarloerra/SocratiCode

    Set up, index, and manage SocratiCode codebase indexing. An agent skill from giancarloerra/SocratiCode.

    3.3k GitHub starsUsed in 1 repo~1.8k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Reference for building on Microsoft Foundry with the azure-ai-projects Python SDK: project clients, versioned agents, evaluations, connections, datasets and indexes.

    3.1k GitHub starsUsed in 6 repos~2.8k tokens
    AI & LLM EngineeringAuto-check passed
  • Fine-Tuning Expert

    Jeffallan/claude-skills

    Guides LLM fine-tuning with LoRA and QLoRA through Hugging Face PEFT, from dataset validation and training checks to adapter merging, quantization and deployment.

    12k GitHub starsUsed in 1 repo~1.7k tokens
    AI & LLM EngineeringAuto-check passed
  • Docs Planner

    strands-agents/harness-sdk

    Identify documentation gaps and prioritize the docs backlog.

    8.7k GitHub stars~821 tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from ai-twinkle/Eval

  • Add Benchmark

    ai-twinkle/Eval

    為 Twinkle Eval 新增一個評測 benchmark(IFEval、BFCL、RAGAS、Text2SQL、Vision MCQ 之類)。涵蓋 CLAUDE.md §6 的完整強制流程:先建 Milestone 與 6 個 Issue、準備 example dataset、實作 Extractor + Scorer 並註冊 PRESETS、與參考框架做分數與速度對比、撰寫…

    117 GitHub stars~1.8k tokensUpdated 23 days ago
    Auto-check passed

Works with

Questions about Run Eval

What does Run Eval do?

用 Twinkle Eval 跑一次評測——建立 config.yaml、下載或指定資料集、驗證設定、執行並讀結果。當使用者說「跑 benchmark」「跑評測」「建 config」「evaluate 這個模型」「twinkle-eval 怎麼跑」「評測結果怎麼看」「為什麼分數是 0」時使用。涵蓋 15 種 evaluationmethod 的 config 差異、27 個內建可下載…. Run Eval is an agent skill from ai-twinkle/Eval.

When should I use Run Eval?

Run Eval fits situations like: AI & LLM Engineering work in your project.

How do I install Run Eval in Claude Code?

Run `npx skills add ai-twinkle/Eval --skill run-eval -a claude-code`. Or copy the skill folder (.claude/skills/run-eval in ai-twinkle/Eval) into .claude/skills/run-eval in your project. Claude Code loads it when a task matches its description.

How do I install Run Eval in Codex?

Run `npx skills add ai-twinkle/Eval --skill run-eval -a codex`. Or copy the skill folder (.claude/skills/run-eval in ai-twinkle/Eval) into .agents/skills/run-eval in your project. Codex loads it when a task matches its description.

Can I use Run Eval in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ai-twinkle/Eval --skill run-eval -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/run-eval, .gemini/skills/run-eval, .github/skills/run-eval and .opencode/skills/run-eval in your project.

What does Run Eval need to run?

Going by SKILL.md and its folder, Run Eval needs the command-line tools its instructions call (jq and git).

Does Run Eval access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Run Eval safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Run Eval use?

Run Eval is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Run Eval use?

About 1.7k tokens (SKILL.md is roughly 6.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.3k tokens, read only when the agent opens those files.

What are the alternatives to Run Eval?

Skills that share tags, products or a category with Run Eval: Chroma Vector Database (Orchestra-Research/AI-Research-SKILLs, 13k stars), CLIP Image-Text Matching (Orchestra-Research/AI-Research-SKILLs, 13k stars), Codebase Management (giancarloerra/SocratiCode, 3.3k stars) and Azure AI Projects Python SDK (microsoft/skills, 3.1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Run Eval?

ai-twinkle (a GitHub organization) maintains it in ai-twinkle/Eval, which has 117 GitHub stars. The repository holds 2 skills in this directory. The repository was last updated on September 15, 2026.

Source: ai-twinkle/Eval on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.