Agent skill

Experiment Design

by voidful in voidful/academic-skills

學術研究實驗設計技能——從研究假設到可重現實驗計畫的完整流程。當使用者需要規劃實驗、設計 ablation study、選擇 baseline、確定評估指標,或問「我應該跑哪些實驗」時,一定要使用此技能。觸發詞包括:實驗設計、experiment design、ablation、baseline、跑什麼實驗、evaluation metric、如何驗證方法。適用於機器學習、NLP、CV…

MITAuto-check passedResearch & Science

Install Experiment Design

skills CLI
$ npx skills add voidful/academic-skills --skill experiment-design -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install voidful/academic-skills experiment-design --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/voidful/academic-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/experiment-design .claude/skills/experiment-design && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
experiment-design
GitHub stars
132
Token cost
~1.2k tokens
SKILL.md length
303 words
Files
6 (incl. references)
Skills in repo
8
Repo updated
First seen
Licence
MIT

At a glance

學術研究實驗設計技能——從研究假設到可重現實驗計畫的完整流程。當使用者需要規劃實驗、設計 ablation study、選擇 baseline、確定評估指標,或問「我應該跑哪些實驗」時,一定要使用此技能。觸發詞包括:實驗設計、experiment design、ablation、baseline、跑什麼實驗、evaluation metric、如何驗證方法。適用於機器學習、NLP、CV…

  • Works in 4 steps: 辨識研究問題:你想回答什麼問題? → 提出核心假設:對問題的預期答案是什麼? → 明確化假設:假設必須具備可測量性與可證偽性 → …
  • Tasks that involve Experimental design
  • SKILL.md covers 概述, 核心設計理念, 實驗設計 Pipeline and 步驟一:研究假設明確化, plus 5 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Experiment Design is an agent skill from voidful/academic-skills. 學術研究實驗設計技能——從研究假設到可重現實驗計畫的完整流程。當使用者需要規劃實驗、設計 ablation study、選擇 baseline、確定評估指標,或問「我應該跑哪些實驗」時,一定要使用此技能。觸發詞包括:實驗設計、experiment design、ablation、baseline、跑什麼實驗、evaluation metric、如何驗證方法。適用於機器學習、NLP、CV 等領域的實驗規劃。

Its SKILL.md is about 1.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files, including reference files (for example `references/ablation-design.md`, `references/baseline-selection.md` and `references/experiment-planning.md`). Compatibility notes: Works with Claude Code, ChatGPT/Codex CLI, and Gemini CLI.

It sits in Research & Science, covering Experimental design and Natural language processing. The repository describes itself as: A complete academic research Skill suite. Supports Claude Code, ChatGPT / Codex CLI, and Gemini CLI. The licence is MIT.

When your agent uses it

  • Tasks that involve Experimental design
  • Tasks that involve Natural language processing

Example prompts

  • “/experiment-design”

Requirements

  • Compatibility (from SKILL.md): Works with Claude Code, ChatGPT/Codex CLI, and Gemini CLI.

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. 辨識研究問題:你想回答什麼問題?
  2. 提出核心假設:對問題的預期答案是什麼?
  3. 明確化假設:假設必須具備可測量性與可證偽性
  4. 分解子假設:將複雜假設拆解為可逐一驗證的子假設

What it can do on your machine

Read from SKILL.md and the folder at commit 71e9c42. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Works with Claude Code, ChatGPT/Codex CLI, and Gemini CLI.

    From compatibility in the SKILL.md frontmatter.

Context cost

Experiment Design loads about 1.2k tokens when it runs, and up to ~9.2k if it reads all its reference files. Until then it costs about 56 tokens; SKILL.md has 303 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~56
When it runs · the whole SKILL.md, loaded when a task matches
~1.2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~9.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from voidful/academic-skills at commit 71e9c42, republished under its MIT licence (© voidful). 303 words, ~1,186 tokens.

Download SKILL.mdSave it as .claude/skills/experiment-design/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
experiment-design
description
學術研究實驗設計技能——從研究假設到可重現實驗計畫的完整流程。當使用者需要規劃實驗、設計 ablation study、選擇 baseline、確定評估指標,或問「我應該跑哪些實驗」時,一定要使用此技能。觸發詞包括:實驗設計、experiment design、ablation、baseline、跑什麼實驗、evaluation metric、如何驗證方法。適用於機器學習、NLP、CV 等領域的實驗規劃。
compatibility
Works with Claude Code, ChatGPT/Codex CLI, and Gemini CLI.
license
MIT
metadata.author
Research Reading Agent
metadata.version
1.0.0

實驗設計技能

概述

本技能提供一套結構化的實驗設計流程,適用於機器學習、自然語言處理、電腦視覺等領域的學術研究。目標是協助研究者從模糊的研究想法出發,產出一份嚴謹、可重現、且具說服力的實驗計畫。

核心設計理念

好的實驗設計應具備以下特質:

  • 可證偽性:實驗結果必須能夠支持或否定研究假設
  • 公平性:所有比較對象在相同條件下評估
  • 可重現性:他人能夠依據描述完整重現實驗
  • 充分性:實驗覆蓋足夠的面向以支撐論文結論

實驗設計 Pipeline

完整的實驗設計遵循以下六步流程:

假設 → 變數 → 指標 → Baseline → Ablation → 計算預算

每一步的產出都是下一步的輸入,形成嚴謹的推導鏈。


步驟一:研究假設明確化

目的

將模糊的研究動機轉化為可驗證的具體假設。

方法
  1. 辨識研究問題:你想回答什麼問題?
  2. 提出核心假設:對問題的預期答案是什麼?
  3. 明確化假設:假設必須具備可測量性與可證偽性
  4. 分解子假設:將複雜假設拆解為可逐一驗證的子假設
假設的品質標準
標準說明
具體性明確指出預期的效果方向與幅度
可測量性可以用量化指標來驗證
可證偽性存在可能否定假設的實驗結果
相關性與研究問題直接相關
範例
  • 不佳:「我們的方法比較好」
  • 良好:「在 SQuAD 2.0 資料集上,加入跨注意力機制後,F1 分數相較於純自注意力基線提升至少 2 個百分點」

詳見:實驗規劃參考


步驟二:變數定義

自變數(Independent Variables)

研究者主動操控的變數,即實驗中「改變的東西」。

  • 模型架構的變體
  • 訓練策略的差異
  • 資料處理方式的不同
依變數(Dependent Variables)

用來衡量實驗結果的變數,即「被測量的東西」。

  • 模型效能指標(準確率、F1、BLEU 等)
  • 效率指標(推論時間、記憶體用量)
  • 品質指標(人工評估分數)
控制變數(Control Variables)

實驗中保持不變的變數,確保比較的公平性。

  • 隨機種子
  • 訓練資料集與切分方式
  • 超參數(非研究對象的部分)
  • 硬體環境
  • 預訓練模型版本
變數控制原則
  1. 單一變數原則:每次實驗僅改變一個自變數
  2. 完整記錄原則:所有變數的值都必須記錄
  3. 合理範圍原則:自變數的取值範圍應有理論依據

詳見:實驗規劃參考


步驟三:評估指標選擇

選擇原則
  1. 領域慣例:優先選擇該領域公認的標準指標
  2. 多面向覆蓋:同時報告效能、效率、穩健性指標
  3. 統計顯著性:報告多次實驗的平均值與標準差
  4. 合理性:指標能真正反映研究假設所關注的面向
常見指標類別
類別指標範例
分類任務Accuracy、Precision、Recall、F1-score、AUC-ROC
生成任務BLEU、ROUGE、METEOR、BERTScore、人工評估
資訊擷取MAP、MRR、NDCG、Recall@K
效率指標FLOPs、參數量、推論延遲、記憶體佔用
穩健性跨資料集表現、對抗樣本準確率
統計檢驗
  • 報告多次隨機種子實驗的平均值與標準差
  • 必要時進行統計顯著性檢驗(如 paired t-test、bootstrap test)
  • 標註統計顯著性水準(p < 0.05, p < 0.01)

步驟四:Baseline 選擇與設定

必選 Baseline 類型
  1. 經典方法:該領域歷史上重要的方法
  2. 當前 SOTA:最新的最佳表現方法
  3. 簡單 Baseline:簡單但合理的基準方法(如隨機、多數類別、TF-IDF)
公平比較原則
  • 使用相同的資料切分
  • 使用相同的評估協定
  • 盡可能使用原作者的程式碼與超參數
  • 若需重新實現,需驗證重現結果與原論文一致
常見錯誤
  • 僅與弱基線比較
  • 未使用最新 SOTA 作為基線
  • 基線的超參數未經調校
  • 比較條件不一致(如不同的預訓練模型)

詳見:Baseline 選擇指南


步驟五:Ablation Study 設計

Ablation Study 是驗證方法中各組件貢獻的關鍵實驗。本技能定義四種 Ablation 模式:

模式一:完整消融(Component Ablation)

逐一移除或替換方法中的各個組件,觀察效能變化。

  • 每次僅移除一個組件
  • 記錄移除後的效能變化
  • 藉此判斷每個組件的貢獻度
模式二:超參敏感度分析(Hyperparameter Sensitivity)

探討關鍵超參數對效能的影響。

  • 選擇 2-4 個最重要的超參數
  • 在合理範圍內變化超參數值
  • 繪製超參數-效能曲線圖
模式三:跨資料集遷移(Cross-Dataset Transfer)

驗證方法的泛化能力。

  • 在多個不同資料集上測試
  • 包含不同規模、不同領域的資料集
  • 分析方法在何種條件下效果最佳或最差
模式四:定性分析(Qualitative Analysis)

透過可視化與案例分析深入理解模型行為。

  • 注意力權重視覺化
  • 成功與失敗案例分析
  • 特徵空間視覺化(如 t-SNE)
  • 錯誤類型分類與統計

詳見:Ablation 設計指南


步驟六:計算資源預估

預估項目
  1. 單次實驗成本

    • GPU 時數
    • 記憶體需求
    • 儲存空間需求
  2. 實驗總量計算

    總 GPU 時數 = 單次時數 × 模型變體數 × 資料集數 × 隨機種子數 × 超參組合數
  3. 安全係數

    • 建議預留 1.5-2 倍的預估資源
    • 考慮除錯、預實驗、追加實驗的需求
資源最佳化策略
  • 先以小規模資料集進行預實驗
  • 使用早停(early stopping)節省訓練時間
  • 善用混合精度訓練(mixed precision)
  • 合理安排實驗優先順序

可重現性要求

實驗計畫必須包含完整的可重現性資訊,確保他人能夠精確重現結果。

必要揭露項目
  1. 硬體環境

    • GPU 型號與數量
    • CPU 規格
    • 記憶體大小
  2. 軟體環境

    • 程式語言版本
    • 深度學習框架版本
    • 關鍵套件版本
  3. 隨機性控制

    • 隨機種子設定
    • 確定性演算法設定
    • 多次實驗的種子列表
  4. 訓練協定

    • 完整的超參數列表
    • 優化器設定
    • 學習率排程
    • 資料增強策略
    • 早停準則
  5. 資料處理

    • 資料集版本與來源
    • 前處理步驟
    • 資料切分方式
  6. 評估協定

    • 評估指標的精確定義
    • 評估頻率
    • 模型選擇準則

詳見:可重現性清單


輸出:結構化實驗計畫文件

本技能的最終產出為一份結構化的實驗計畫文件,包含以下章節:

  1. 研究假設與子假設
  2. 變數定義表
  3. 評估指標與統計方法
  4. Baseline 列表與設定
  5. Ablation Study 設計矩陣
  6. 計算資源預估與時程規劃
  7. 可重現性資訊

使用模板:實驗計畫模板


使用流程

輸入
  • 研究主題或論文草稿
  • 提出的方法描述
  • 可用的計算資源
處理
  1. 引導使用者明確化研究假設
  2. 協助定義自變數、依變數、控制變數
  3. 根據任務類型建議評估指標
  4. 根據研究領域建議 Baseline
  5. 設計 Ablation Study 方案
  6. 預估計算資源需求
輸出
  • 完整的實驗計畫文件(依照模板格式)
  • 實驗優先順序建議
  • 潛在風險與應對方案

品質檢查清單

在完成實驗計畫後,請確認以下項目:

  • 每個研究假設都有對應的實驗來驗證
  • 所有自變數的取值範圍已明確定義
  • 控制變數已完整列出
  • 評估指標涵蓋多個面向
  • Baseline 包含經典方法、SOTA、簡單基線
  • Ablation Study 覆蓋所有提出的組件
  • 計算資源預估合理且包含安全係數
  • 可重現性資訊完整
  • 統計檢驗方法已確定

參考資源

© voidful, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files (references) in experiment-design of voidful/academic-skills.

  • SKILL.md
  • references/ablation-design.md
  • references/baseline-selection.md
  • references/experiment-planning.md
  • references/reproducibility-checklist.md
  • templates/experiment-plan.md

Open the folder on GitHubat commit 71e9c42

Compare with similar skills

Experiment Design next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Experiment Design compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Experiment Design this skillvoidful/academic-skills132—~1.2kAutomated safety check: PassMIT
Theory Variable Method MatrixDrchronx/ai-agent-research-starter-kit134—~741Automated safety check: PassCustom licence
Scientific Critical Thinkingweapp-tailwindcss/weapp-tailwindcss1.9k23 repos~5.9kAutomated safety check: NotesMIT
Claim-Driven Experiment PlannerzjYao36/Auto-Research-Refine1287 repos~2.3kAutomated safety check: NotesNone
Benchmark Paper TemplateHKUSTDial/Supervisor-Skills8.4k—~2.8kAutomated safety check: PassCC-BY-4.0
Research Refine PipelinezjYao36/Auto-Research-Refine1286 repos~1.4kAutomated safety check: NotesNone

Similar skills

  • Theory Variable Method Matrix

    Drchronx/ai-agent-research-starter-kit

    Build and maintain a theory-variable-method matrix from papers, Zotero metadata, Obsidian paper cards, and research notes.

    134 GitHub stars~741 tokensUpdated 4 mo ago
    Research & ScienceAuto-check passed
  • Scientific Critical Thinking

    weapp-tailwindcss/weapp-tailwindcss

    Evaluate research rigor. An agent skill from weapp-tailwindcss/weapp-tailwindcss.

    1.9k GitHub starsUsed in 23 repos~5.9k tokens
    Research & ScienceAuto-check: notes
  • Claim-Driven Experiment Planner

    zjYao36/Auto-Research-Refine

    Turns a refined research proposal into a claim-to-evidence-to-run-order roadmap instead of a sprawling benchmark wishlist.

    128 GitHub starsUsed in 7 repos~2.3k tokens
    Research & ScienceAuto-check: notes
  • Benchmark Paper Template

    HKUSTDial/Supervisor-Skills

    Structures benchmark and evaluation papers around five pillars, with a completeness audit, an Introduction logic chain, a section skeleton and a pre-submission checklist.

    8.4k GitHub stars~2.8k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Research Refine Pipeline

    zjYao36/Auto-Research-Refine

    Chains research-refine and experiment-plan to turn a vague research direction into a focused proposal and a claim-driven experiment roadmap.

    128 GitHub starsUsed in 6 repos~1.4k tokens
    Research & ScienceAuto-check: notes
  • Metabolic Study Planner

    aiming-lab/AutoResearchClaw

    Turns a broad metabolic modelling topic into a concrete, paper-shaped plan with organism, model, perturbations, metrics and figures before any FBA code is written.

    15k GitHub stars~1.9k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed

More from voidful/academic-skills

All 8 skills in this repo
  • Academic Research

    voidful/academic-skills

    Complete academic research skill suite covering the full pipeline: paper reading (read/explain papers with storytelling), idea generation (brainstorm research directions), experiment design (plan…

    132 GitHub stars~887 tokensUpdated 6 mo ago
    Auto-check passed
  • Idea Generation

    voidful/academic-skills

    學術研究的 Idea 產生技能——從發散到收斂,系統化地產出高品質研究構想。當使用者想腦力激盪研究方向、找新 research idea、或問「我接下來可以做什麼研究」時,一定要使用此技能。觸發詞包括:brainstorm、想 idea、研究方向、下一步做什麼、有什麼可以研究的、找 gap、research proposal。適用於任何階段的學術研究構想生成。

    132 GitHub stars~1.6k tokensUpdated 6 mo ago
    Auto-check passed
  • Paper Reading

    voidful/academic-skills

    太奶讀論文 — 一位百歲阿嬤用繁體中文、生活比喻和動漫梗,帶你讀懂學術論文。當使用者提供論文 PDF、arXiv 連結、或貼上論文文字,並想理解論文內容時,一定要使用此技能。觸發詞包括:讀論文、解釋論文、看不懂、幫我理解這篇、這篇在說什麼、paper reading、explain this paper。適用於任何學術論文的直觀導讀。

    132 GitHub stars~1.4k tokensUpdated 6 mo ago
    Auto-check passed
  • Paper Review

    voidful/academic-skills

    學術論文審稿技能 — 以結構化四步驟流程完成深度論文審查,涵蓋批判性審查、分數預測、要點精煉與正式審稿產出。當使用者需要 review 一篇論文、模擬 reviewer 反應、評估論文能否被接收、或幫助判斷論文優缺點時,一定要使用此技能。觸發詞包括:review 這篇、幫我審稿、reviewer 會怎麼說、這篇能上嗎、paper review、給分數、找…

    132 GitHub stars~2k tokensUpdated 6 mo ago
    Auto-check passed
  • Paper Writing

    voidful/academic-skills

    頂級會議論文寫作技能——以嚴格 reviewer 視角指導從草稿到終稿的完整寫作流程。當使用者要寫論文、改善論文草稿、修改特定章節(introduction、method、experiments、conclusion)、潤色學術英文、回應 reviewer 意見,或問「這段怎麼寫」時,一定要使用此技能。觸發詞包括:寫論文、paper writing、improve my…

    132 GitHub stars~1.9k tokensUpdated 6 mo ago
    Auto-check passed
  • Proof Writer

    voidful/academic-skills

    數學證明撰寫技能 — 從主張提取到 LaTeX 排版的完整證明工作流。當使用者需要撰寫或驗證數學定理、引理、命題的形式證明,或需要推導公式、整理理論分析時,一定要使用此技能。觸發詞包括:數學證明、prove、theorem、lemma、proposition、推導、理論分析、寫 proof、LaTeX 數學、formal…

    132 GitHub stars~1.7k tokensUpdated 6 mo ago
    Auto-check passed

Questions about Experiment Design

What does Experiment Design do?

學術研究實驗設計技能——從研究假設到可重現實驗計畫的完整流程。當使用者需要規劃實驗、設計 ablation study、選擇 baseline、確定評估指標,或問「我應該跑哪些實驗」時,一定要使用此技能。觸發詞包括:實驗設計、experiment design、ablation、baseline、跑什麼實驗、evaluation metric、如何驗證方法。適用於機器學習、NLP、CV…. Experiment Design is an agent skill from voidful/academic-skills.

When should I use Experiment Design?

Experiment Design fits situations like: tasks that involve Experimental design; tasks that involve Natural language processing.

How do I install Experiment Design in Claude Code?

Run `npx skills add voidful/academic-skills --skill experiment-design -a claude-code`. Or copy the skill folder (experiment-design in voidful/academic-skills) into .claude/skills/experiment-design in your project. Claude Code loads it when a task matches its description.

How do I install Experiment Design in Codex?

Run `npx skills add voidful/academic-skills --skill experiment-design -a codex`. Or copy the skill folder (experiment-design in voidful/academic-skills) into .agents/skills/experiment-design in your project. Codex loads it when a task matches its description.

Can I use Experiment Design in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add voidful/academic-skills --skill experiment-design -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/experiment-design, .gemini/skills/experiment-design, .github/skills/experiment-design and .opencode/skills/experiment-design in your project.

What does Experiment Design need to run?

SKILL.md names no scripts, command-line tools or credentials: Experiment Design is instructions for the agent only. Compatibility (from SKILL.md): Works with Claude Code, ChatGPT/Codex CLI, and Gemini CLI..

Does Experiment Design access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Experiment Design safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Experiment Design use?

Experiment Design is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Experiment Design use?

About 1.2k tokens (SKILL.md is roughly 4.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 8k tokens, read only when the agent opens those files.

What are the alternatives to Experiment Design?

Skills that share tags, products or a category with Experiment Design: Theory Variable Method Matrix (Drchronx/ai-agent-research-starter-kit, 134 stars), Scientific Critical Thinking (weapp-tailwindcss/weapp-tailwindcss, 1.9k stars), Claim-Driven Experiment Planner (zjYao36/Auto-Research-Refine, 128 stars) and Benchmark Paper Template (HKUSTDial/Supervisor-Skills, 8.4k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Experiment Design?

voidful (a GitHub user) maintains it in voidful/academic-skills, which has 132 GitHub stars. The repository holds 8 skills in this directory. The repository was last updated on April 4, 2026.

Source: voidful/academic-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.