Chroma Vector Database
Orchestra-Research/AI-Research-SKILLs
Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.
用 Twinkle Eval 跑一次評測——建立 config.yaml、下載或指定資料集、驗證設定、執行並讀結果。當使用者說「跑 benchmark」「跑評測」「建 config」「evaluate 這個模型」「twinkle-eval 怎麼跑」「評測結果怎麼看」「為什麼分數是 0」時使用。涵蓋 15 種 evaluationmethod 的 config 差異、27 個內建可下載…
$ npx skills add ai-twinkle/Eval --skill run-eval -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install ai-twinkle/Eval run-eval --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/ai-twinkle/Eval.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/run-eval .claude/skills/run-eval && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "run-eval" agent skill from https://github.com/ai-twinkle/Eval/tree/main/.claude/skills/run-eval into .claude/skills/run-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "run-eval", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/ai-twinkle/Eval/tree/main/.claude/skills/run-evalType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add ai-twinkle/Eval --skill run-eval -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install ai-twinkle/Eval run-eval --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ai-twinkle/Eval.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills/run-eval .agents/skills/run-eval && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "run-eval" agent skill from https://github.com/ai-twinkle/Eval/tree/main/.claude/skills/run-eval into .agents/skills/run-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "run-eval", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ai-twinkle/Eval --skill run-eval -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install ai-twinkle/Eval run-eval --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ai-twinkle/Eval.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills/run-eval .cursor/skills/run-eval && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "run-eval" agent skill from https://github.com/ai-twinkle/Eval/tree/main/.claude/skills/run-eval into .cursor/skills/run-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "run-eval", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/ai-twinkle/Eval.git --path .claude/skills/run-eval--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add ai-twinkle/Eval --skill run-eval -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install ai-twinkle/Eval run-eval --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ai-twinkle/Eval.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills/run-eval .gemini/skills/run-eval && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "run-eval" agent skill from https://github.com/ai-twinkle/Eval/tree/main/.claude/skills/run-eval into .gemini/skills/run-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "run-eval", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install ai-twinkle/Eval run-evalInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add ai-twinkle/Eval --skill run-eval -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/ai-twinkle/Eval.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills/run-eval .github/skills/run-eval && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "run-eval" agent skill from https://github.com/ai-twinkle/Eval/tree/main/.claude/skills/run-eval into .github/skills/run-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "run-eval", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ai-twinkle/Eval --skill run-eval -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install ai-twinkle/Eval run-eval --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ai-twinkle/Eval.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills/run-eval .opencode/skills/run-eval && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "run-eval" agent skill from https://github.com/ai-twinkle/Eval/tree/main/.claude/skills/run-eval into .opencode/skills/run-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "run-eval", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
run-eval用 Twinkle Eval 跑一次評測——建立 config.yaml、下載或指定資料集、驗證設定、執行並讀結果。當使用者說「跑 benchmark」「跑評測」「建 config」「evaluate 這個模型」「twinkle-eval 怎麼跑」「評測結果怎麼看」「為什麼分數是 0」時使用。涵蓋 15 種 evaluationmethod 的 config 差異、27 個內建可下載…
Run Eval is an agent skill from ai-twinkle/Eval. 用 Twinkle Eval 跑一次評測——建立 config.yaml、下載或指定資料集、驗證設定、執行並讀結果。當使用者說「跑 benchmark」「跑評測」「建 config」「evaluate 這個模型」「twinkle-eval 怎麼跑」「評測結果怎麼看」「為什麼分數是 0」時使用。涵蓋 15 種 evaluationmethod 的 config 差異、27 個內建可下載 benchmark、--validate / --dry-run / --resume,以及用 unparsedrate 診斷 extractor 失效。
Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/config-reference.md`).
It sits in AI & LLM Engineering. It works with OpenAI. The repository describes itself as: High-performance LLM evaluation framework with parallel API calls — up to 17× faster than sequential tools. Supports box, math, and logit-based evaluation. The licence is MIT.
5 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 608273c. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
jqgitFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Run Eval loads about 1.7k tokens when it runs, and up to ~4k if it reads all its reference files. Until then it costs about 71 tokens; SKILL.md has 342 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from ai-twinkle/Eval at commit 608273c, republished under its MIT licence (© ai-twinkle). 342 words, ~1,667 tokens.
.claude/skills/run-eval/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.Twinkle Eval 只做一件事——拿題目去呼叫已經在外部運行的 OpenAI 相容端點。
模型要自己先起好(vLLM、Ollama、OpenAI、NVIDIA Build 都行),再把 base_url 填進 config。
端點沒回應時,本專案的正確行為是依 max_retries 重試後報錯退出,不會也不該嘗試
重啟服務(CLAUDE.md 原則 G)。
twinkle-eval --init # 列出全部 11 個範本
twinkle-eval --init multiple_choice # 產生 configs/multiple_choice.yaml
twinkle-eval --init all # 全部產生到 configs/範本存放在 twinkle_eval/templates/,--init 直接掃該目錄。
llm_api:
base_url: "http://localhost:8000/v1" # 必填
api_key: "EMPTY" # 必填(本地 vLLM 隨便填)
api_rate_limit: -1 # QPS,-1 為不限
max_retries: 3
timeout: 600
disable_ssl_verify: false
model:
name: "my-model" # 必填,會寫進結果路徑與紀錄
temperature: 0.0
top_p: 0.9
max_tokens: 4096
extra_body: # 傳給 API 的額外參數
evaluation:
dataset_paths: # 必填,list(即使只有一個)
- "datasets/example/tmmluplus/"
evaluation_method: "box" # 必填,見下表
repeat_runs: 1 # >1 時算平均與標準差
shuffle_options: false # 選項隨機排列
logging:
level: "INFO"每個 evaluation_method 的必填欄位與 strategy_config 參數不同,
完整對照見 references/config-reference.md。
| 方法 | 用在 | 備註 |
|---|---|---|
pattern | 選擇題,通用首選 | 正則比對,含中英文預設模式 |
box | 選擇題,推理模型 | 提取 \boxed{} / \box{};需要 system_prompt |
logit | 多選項題 | 比較各選項 log-probability,不依賴輸出格式;選項數不限 |
math | 數學推理 | \boxed{} + MathRuler;需 [math] |
regex_match | BBH 之類自由格式 | ⚠️ system_prompt 不生效,見下方說明 |
custom_regex | 自訂格式 | 必須設 strategy_config.patterns |
ifeval / ifbench | 指令遵循 | 需 [ifeval] / [ifbench] + nltk 資料 |
bfcl_fc / bfcl_prompt | 函式呼叫 | FC 走 tools API,prompt 走注入 |
niah | 長文本大海撈針 | |
ragas | RAG 品質 | |
text2sql | Text-to-SQL | 需設 text2sql_db_base_path |
asr | 語音辨識 | 需 [asr];llm_api.type: whisper 或多模態 |
vision_mcq | 視覺多選題 | 需 [vision](縮放用) |
⚠️
evaluation.system_prompt只對box與math生效。models/openai.py的_build_messages()以白名單決定是否送出 system message (method in {"box", "math"}),其他方法即使在 config 設了system_prompt也不會進 request。regex_match這類需要指定輸出格式的方法,格式要求必須寫進資料集的question欄位裡。 (專案的templates/regex_match.yaml目前有同樣的誤導,見 #144)
twinkle-eval --list-strategies # 執行期確認可用方法twinkle-eval --download-dataset list # 列出 27 個內建 benchmark
twinkle-eval --download-dataset mmlu # 短名稱
twinkle-eval --download-dataset tmmluplus gsm8k
twinkle-eval --download-dataset all內建短名稱:mmlu mmlu_pro mmlu_redux tmmluplus supergpqa gpqa formosa_bench
gsm8k aime2025 bbh ifeval ifbench bfcl needlebench longbench wikieval
librispeech aishell1 fleurs common_voice mmbench mmstar mmmu pope
spider bird spider2_lite
gpqa 是 gated dataset,會互動式要求 HuggingFace token。
先用 example 資料集驗證流程跑得通再下載完整 benchmark:
datasets/example/ 底下每個 benchmark 都有 10–30 筆的子集,跑一次只要幾秒。
含真實 API 金鑰的 config 絕對不得 commit(CLAUDE.md 原則 E)。這三種前綴/後綴已被
.gitignore 全局排除,本機測試一律用其中之一:
config_local_*.yaml
config_test_*.yaml
*.local.yaml寫入任何含金鑰的 config 前,先確認該路徑已在 .gitignore 中。不確定就檢查:
git check-ignore -v config_local_myrun.yaml # 有輸出 = 已被忽略
git diff --staged | grep -i "api_key" # commit 前確認twinkle-eval --validate --config config_local_myrun.yaml # 只驗設定與資料集,不呼叫 API
twinkle-eval --dry-run --config config_local_myrun.yaml # 顯示評測計畫,不呼叫 API
twinkle-eval --config config_local_myrun.yaml # 正式跑
twinkle-eval --config config_local_myrun.yaml --export json csv html excel永遠先跑 --validate 再跑正式評測。 資料集路徑錯、缺必填欄位、格式不符都會在這一步
抓到,省下一整輪 API 費用與時間。
中斷後續跑(⚠️ 目前失效,見 #145):
twinkle-eval --resume 20260825_1430 --config config_local_myrun.yamlresults/
├── results_{timestamp}.json # 整體摘要
└── eval_results_{timestamp}_run{N}.jsonl # 各題明細(append 模式)摘要含 dataset_results(各資料集的 average_accuracy、average_std、
average_pass_at_k、total_unparsed_count)與 duration_seconds。
config 欄位是移除 api_key 後的設定。
明細每行含 question_id / sample_id / question / correct_answer /
predicted_answer / is_correct / llm_output / llm_reasoning_output / token 用量。
各路徑另有增補欄位(logit 的 logprob_scores、ifeval 的四個指標、vision_mcq 的
image_path、asr 的 wer / cer)。
⚠️ 實際輸出沒有
file與timestamp欄位(CLAUDE.md §10 的規範與實作不符)。 這也讓--resume目前無法運作:它以{file}|{question_id}作為已完成紀錄的 key,file缺失使得 key 變成|{idx},與比對端的{檔案路徑}|永遠不匹配, 結果是一題都不會跳過、整輪重跑並 append 出重複列。見 #145。
# 快速看正確率
jq -r '.dataset_results | to_entries[] | "\(.key): \(.value.average_accuracy)"' \
results/results_*.json
# 撈出所有答錯的題目
jq -c 'select(.is_correct == false) | {question_id, predicted_answer, correct_answer}' \
results/eval_results_*_run0.jsonl | head先看 unparsed_rate——這是判斷「模型答錯」還是「extractor 沒抓到」的關鍵。
# 資料集層級(摘要 JSON 用 average_unparsed_rate,per-file 才叫 unparsed_rate)
jq -r '.dataset_results | to_entries[] | "\(.key): \(.value.average_unparsed_rate)"' \
results/results_*.json| 症狀 | 多半是 |
|---|---|
unparsed_rate 高(>10%) | extractor 沒對上輸出格式 |
unparsed_rate ≈ 0 但分數低 | 模型是真的答錯 |
| 全部 0 分且無 unparsed | ground truth 欄位或正規化對不上 |
| 「所有資料集評測均失敗」 | 資料集路徑、格式,或 API 端點問題 |
extractor 沒抓到時,撈幾筆 llm_output 出來看實際格式:
jq -r 'select(.predicted_answer == null) | .llm_output' \
results/eval_results_*_run0.jsonl | head -3常見成因:
box 方法但沒設 system_prompt → 模型不知道要用 \boxed{},自然抓不到<think>...</think> / <reason> /
<reasoning> 標籤對;只有結尾 tag 而無開頭 tag 會被視為格式不合格而原樣保留content 為 null(vLLM skip_special_tokens=true)→ 會回退讀 reasoning,
再回退 reasoning_content(vLLM 0.18+ 改名,兩者都支援)並行度由 ThreadPoolExecutor 預設值決定,用 llm_api.api_rate_limit 節流:
-1(不限)repeat_runs > 1 會線性放大時間,但能得到標準差——量化模型穩定性時才開。
© ai-twinkle, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 1 other file (references) in .claude/skills/run-eval of ai-twinkle/Eval.
Open the folder on GitHubat commit 608273c
Run Eval next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Run Eval this skillai-twinkle/Eval | 117 | — | ~1.7k | Automated safety check: Pass | MIT | |
| Chroma Vector DatabaseOrchestra-Research/AI-Research-SKILLs | 13k | 8 repos | ~2.3k | Automated safety check: Pass | MIT | |
| CLIP Image-Text MatchingOrchestra-Research/AI-Research-SKILLs | 13k | 8 repos | ~1.7k | Automated safety check: Pass | MIT | |
| Codebase Managementgiancarloerra/SocratiCode | 3.3k | 1 repos | ~1.8k | Automated safety check: Pass | AGPL-3.0 | |
| Azure AI Projects Python SDKmicrosoft/skills | 3.1k | 6 repos | ~2.8k | Automated safety check: Pass | MIT | |
| Fine-Tuning ExpertJeffallan/claude-skills | 12k | 1 repos | ~1.7k | Automated safety check: Pass | MIT |
Orchestra-Research/AI-Research-SKILLs
Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.
Orchestra-Research/AI-Research-SKILLs
Explains OpenAI's CLIP model for zero-shot image classification, image-text similarity, semantic image search and content moderation, with install steps and code patterns.
giancarloerra/SocratiCode
Set up, index, and manage SocratiCode codebase indexing. An agent skill from giancarloerra/SocratiCode.
microsoft/skills
Reference for building on Microsoft Foundry with the azure-ai-projects Python SDK: project clients, versioned agents, evaluations, connections, datasets and indexes.
Jeffallan/claude-skills
Guides LLM fine-tuning with LoRA and QLoRA through Hugging Face PEFT, from dataset validation and training checks to adapter merging, quantization and deployment.
strands-agents/harness-sdk
Identify documentation gaps and prioritize the docs backlog.
ai-twinkle/Eval
為 Twinkle Eval 新增一個評測 benchmark(IFEval、BFCL、RAGAS、Text2SQL、Vision MCQ 之類)。涵蓋 CLAUDE.md §6 的完整強制流程:先建 Milestone 與 6 個 Issue、準備 example dataset、實作 Extractor + Scorer 並註冊 PRESETS、與參考框架做分數與速度對比、撰寫…
Works with
Categories
用 Twinkle Eval 跑一次評測——建立 config.yaml、下載或指定資料集、驗證設定、執行並讀結果。當使用者說「跑 benchmark」「跑評測」「建 config」「evaluate 這個模型」「twinkle-eval 怎麼跑」「評測結果怎麼看」「為什麼分數是 0」時使用。涵蓋 15 種 evaluationmethod 的 config 差異、27 個內建可下載…. Run Eval is an agent skill from ai-twinkle/Eval.
Run Eval fits situations like: AI & LLM Engineering work in your project.
Run `npx skills add ai-twinkle/Eval --skill run-eval -a claude-code`. Or copy the skill folder (.claude/skills/run-eval in ai-twinkle/Eval) into .claude/skills/run-eval in your project. Claude Code loads it when a task matches its description.
Run `npx skills add ai-twinkle/Eval --skill run-eval -a codex`. Or copy the skill folder (.claude/skills/run-eval in ai-twinkle/Eval) into .agents/skills/run-eval in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ai-twinkle/Eval --skill run-eval -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/run-eval, .gemini/skills/run-eval, .github/skills/run-eval and .opencode/skills/run-eval in your project.
Going by SKILL.md and its folder, Run Eval needs the command-line tools its instructions call (jq and git).
SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Run Eval is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.7k tokens (SKILL.md is roughly 6.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.3k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Run Eval: Chroma Vector Database (Orchestra-Research/AI-Research-SKILLs, 13k stars), CLIP Image-Text Matching (Orchestra-Research/AI-Research-SKILLs, 13k stars), Codebase Management (giancarloerra/SocratiCode, 3.3k stars) and Azure AI Projects Python SDK (microsoft/skills, 3.1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
ai-twinkle (a GitHub organization) maintains it in ai-twinkle/Eval, which has 117 GitHub stars. The repository holds 2 skills in this directory. The repository was last updated on September 15, 2026.
Source: ai-twinkle/Eval on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.