Agent skill

Experiment Design

by flonat in flonat/flonat-research

Design empirical studies through power analysis, pre-analysis planning, QSF parsing, and survey architecture.

MITAuto-check passedResearch & Science

Install Experiment Design

skills CLI
$ npx skills add flonat/flonat-research --skill experiment-design -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install flonat/flonat-research experiment-design --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/flonat/flonat-research.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/experiment-design .claude/skills/experiment-design && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
experiment-design
GitHub stars
146
Token cost
~2.3k tokens
SKILL.md length
973 words
Files
7 (incl. references)
Skills in repo
83
Repo updated
First seen
Licence
MIT

At a glance

Design empirical studies through power analysis, pre-analysis planning, QSF parsing, and survey architecture.

  • Works in 4 steps: Interview — ask for → Generate script — R (DeclareDesign/pwr)… → Execute and report — produce a sample… → …
  • Specifying sampling
  • SKILL.md covers Modes, When to Use, When NOT to Use and Shared References, plus 5 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Experiment Design is an agent skill from flonat/flonat-research. Design empirical studies through power analysis, pre-analysis planning, QSF parsing, and survey architecture. Use when specifying sampling, measurement, treatment, or analysis before data collection. Not for causal identification alone; use $causal-design.

Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files, including reference files (for example `references/identification-strategies.md`, `references/known-scales-registry.md` and `references/pap-template.md`).

It sits in Research & Science, covering Experimental design. The repository describes itself as: Shareable Claude Code + Codex infrastructure for PhD researchers — skills, agents, hooks, and rules for academic workflows. The licence is MIT.

When your agent uses it

  • Specifying sampling
  • Analysis before data collection

Example prompts

  • “/experiment-design”

Requirements

  • Python 3
  • Pre-approved tools (allowed-tools): Bash(uv*, Rscript*, R*, mkdir*, ls*), Read, Write, Edit, Glob, Grep, AskUserQuestion

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Interview — ask for
  2. Generate script — R (DeclareDesign/pwr) or Python (statsmodels.stats.power)
  3. Execute and report — produce a sample size table showing N for power = {0.80, 0.90, 0.95}
  4. Write to project — save script to code/power_analysis.R (or .py), results to output/power_analysis_results.md

What it can do on your machine

Read from SKILL.md and the folder at commit da27600. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Bash(uv*
    • Rscript*
    • R*
    • mkdir*
    • ls*)
    • Read
    • Write
    • Edit
    • Glob
    • Grep

    …and 1 more on the same allowed-tools line.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Experiment Design loads about 2.3k tokens when it runs, and up to ~9.2k if it reads all its reference files. Until then it costs about 69 tokens; SKILL.md has 973 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~69
When it runs · the whole SKILL.md, loaded when a task matches
~2.3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~9.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from flonat/flonat-research at commit da27600, republished under its MIT licence (© flonat). 973 words, ~2,292 tokens.

Download SKILL.mdSave it as .claude/skills/experiment-design/SKILL.md (or your agent's skills folder). This skill also uses 6 other files; get the full folder from GitHub.
name
experiment-design
description
Design empirical studies through power analysis, pre-analysis planning, QSF parsing, and survey architecture. Use when specifying sampling, measurement, treatment, or analysis before data collection. Not for causal identification alone; use $causal-design.
allowed-tools
Bash(uv*, Rscript*, R*, mkdir*, ls*), Read, Write, Edit, Glob, Grep, AskUserQuestion
argument-hint
[--mode power|design|pap|survey] [qsf-file or project-path]

Experiment Design

Interview-driven design workflow producing design documents, power analysis scripts, and pre-analysis plans.

Modes

ModeWhat it producesEntry point
PowerPower analysis script + sample size table"How many participants do I need?"
DesignFull design document (hypotheses, conditions, measures, randomization)"Design my experiment"
PAPPre-analysis plan (AEA/OSF/EGAP format)"Write a PAP"
SurveyStructured survey specification from natural language or QSF"Build a survey" / "Parse my Qualtrics"

Default: Design. If user provides a .qsf file, auto-select Survey mode.

When to Use

  • Designing a new experiment or survey
  • Calculating required sample sizes
  • Writing or auditing a pre-analysis plan
  • Parsing a Qualtrics .qsf file to understand its structure
  • Building a survey specification from a natural language description

When NOT to Use

  • Running the analysis → data-analysis
  • Auditing identification strategy for observational studies → causal-design
  • Generating synthetic test data → synthetic-data

Shared References

  • Method probing questions: shared/method-probing-questions.md — ask before designing (Experiments/RCTs, Survey sections)
  • Validation tiers: shared/validation-tiers.md — tier determines required power and pre-registration
  • Escalation protocol: shared/escalation-protocol.md — escalate when design has validity threats
  • Engagement-stratified sampling: shared/engagement-stratified-sampling.md — stratify social media samples by engagement
  • Inter-coder reliability: shared/intercoder-reliability.md — reliability planning for content analysis designs

Mode: Power

Read references/power-analysis-recipes.md for language-specific code patterns.

Workflow
  1. Interview — ask for:
    • Primary outcome variable and expected effect size (or domain norms)
    • Design type (between-subjects, within-subjects, factorial, cluster-randomized)
    • Number of conditions/groups
    • Significance level (default: 0.05) and desired power (default: 0.80)
    • Any clustering or stratification
  2. Generate script — R (DeclareDesign/pwr) or Python (statsmodels.stats.power)
  3. Execute and report — produce a sample size table showing N for power = {0.80, 0.90, 0.95}
  4. Write to project — save script to code/power_analysis.R (or .py), results to output/power_analysis_results.md

HPC escalation: If the power analysis uses Monte Carlo simulation (e.g., DeclareDesign with >10k replications, or a multi-design sweep), move execution to [HPC cluster] — drop the simulation script into hpc/ with templates/slurm/array.sbatch (array over seeds/designs). The SHA-logging snippet in the template pins results to the DGP version. See docs/guides/hpc.md.

Effect Size Guidance

If the user doesn't know the expected effect size, guide them:

SourceHow to use
Prior literature"What did similar studies find?"
Pilot dataCalculate from pilot descriptives
SESOI"What's the smallest effect worth detecting?"
Domain normsCohen's benchmarks as absolute last resort (small=0.2, medium=0.5, large=0.8 for d)

Never default to Cohen's benchmarks without acknowledging they are arbitrary.


Mode: Design

Workflow
  1. Research question interview — structured questions:
    • What is the causal question?
    • What is the treatment / intervention?
    • What is the primary outcome? Secondary outcomes?
    • What is the target population?
    • What is the assignment mechanism? (random, stratified, clustered, matched)
  2. Design specification — produce a structured document covering:
    • Hypotheses (directional, with expected signs)
    • Conditions (treatment arms, control)
    • Randomization procedure
    • Outcome measures and scales
    • Sample and recruitment strategy
    • Timeline
  3. Identification check — state the estimand, identifying assumptions, and potential threats
  4. Write design document — save to docs/experiment-design.md or project-appropriate location
  5. Lock the design — record in MEMORY.md Estimand Registry (if project has one) or flag for the user to lock before analysis

This design document is what data-analysis Phase 3 checks for before allowing estimation.


Mode: PAP

Read references/pap-template.md for the full template structure.

Workflow
  1. Check for existing design — look for design document in docs/, log/plans/, or .context/
  2. If no design exists — run Design mode first, then continue
  3. Generate PAP — structured pre-analysis plan following AEA/OSF/EGAP conventions:
    • Study information (title, authors, IRB, registration)
    • Design overview (hypotheses, conditions, randomization, sample)
    • Data collection (instruments, timing, blinding)
    • Analysis plan (main specification, inference, multiple testing corrections)
    • Robustness and sensitivity analyses
    • Power analysis (reference or embed results from Power mode)
  4. Output — save to docs/pre-analysis-plan.md (or .tex if user prefers LaTeX)
  5. Registration prompt — remind user to register at AEA RCT Registry, OSF, or EGAP

Show full SKILL.md (366 more words)Show less

Mode: Survey

Survey mode is a first-class capability with two entry points: QSF parsing and natural language construction.

Entry Point A: QSF Parsing

When user provides a Qualtrics .qsf file:

  1. Parse the JSON — extract survey flow, blocks, questions, embedded data, skip logic
  2. Map question types — read references/qsf-parsing-guide.md for type mapping:
    • Matrix, Likert, slider, numeric, constant sum, rank order, best-worst scaling
    • Text entry (single-line, multi-line, essay)
    • Multiple choice (single answer, multiple answer)
  3. Detect design elements:
    • Factorial conditions from randomizer blocks and embedded data
    • Attention checks and comprehension checks
    • Known validated scales (read references/known-scales-registry.md)
    • Skip/display logic and branching
  4. Produce structured summary — questions, conditions, scales, logic flow, warnings
  5. Flag issues — missing attention checks, unbalanced conditions, potential order effects
Entry Point B: Natural Language Construction

When user describes an experiment in natural language:

  1. Parse the design — extract factorial structure from text
    • Example: "3 (Source: AI vs Human vs None) x 2 (Product: Hedonic vs Utilitarian)" → 3x2 between-subjects
  2. Interview for details:
    • What DVs to measure? (recommend appropriate scale types)
    • What manipulation checks?
    • What attention/comprehension checks?
    • Demographics and control variables?
  3. Generate survey specification — structured document with:
    • Survey flow (consent → demographics → manipulation → DVs → manipulation check → debrief)
    • Question text and response options for each item
    • Condition assignments and randomization logic
    • Attention check placement (read references/survey-design-checklist.md)
  4. Scale recommendations — read references/known-scales-registry.md for validated scales matching the constructs
Survey Quality Checks

Read references/survey-design-checklist.md for the full checklist. Key checks:

  • Attention checks: At least 1 per 5 minutes of survey length. Calibrated pass rates (85-95%). See Krosnick (1991), Meade & Craig (2012).
  • Response style mitigation: Flag scales vulnerable to acquiescence bias. Recommend reverse-coding for scales with 4+ items in the same direction.
  • Scale quality: Semantic type detection (satisfaction, trust, intention, risk). Reverse-coded item flagging. Well-known scale recognition.
  • Order effects: Randomize question order within blocks where appropriate. Flag potential priming from early questions.

Cross-References

ResourceWhen read
references/power-analysis-recipes.mdPower mode
references/pap-template.mdPAP mode
references/survey-design-checklist.mdSurvey mode (quality checks)
references/identification-strategies.mdDesign mode (identification check)
references/qsf-parsing-guide.mdSurvey mode (QSF parsing)
references/known-scales-registry.mdSurvey mode (scale recognition)
design-before-results ruleDesign + PAP modes produce the locked design
data-analysis skillConsumes the design document
causal-design skillFor observational identification (not experiments)
synthetic-data skillFor generating pilot data

© flonat, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 6 other files (references) in skills/experiment-design of flonat/flonat-research.

  • SKILL.md
  • references/identification-strategies.md
  • references/known-scales-registry.md
  • references/pap-template.md
  • references/power-analysis-recipes.md
  • references/qsf-parsing-guide.md
  • references/survey-design-checklist.md

Open the folder on GitHubat commit da27600

Compare with similar skills

Experiment Design next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Experiment Design compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Experiment Design this skillflonat/flonat-research146—~2.3kAutomated safety check: PassMIT
Scientific Critical Thinkingweapp-tailwindcss/weapp-tailwindcss1.9k22 repos~5.9kAutomated safety check: NotesMIT
Benchmark Paper TemplateHKUSTDial/Supervisor-Skills8.7k—~2.8kAutomated safety check: PassCC-BY-4.0
Claim-Driven Experiment PlannerzjYao36/Auto-Research-Refine1286 repos~2.3kAutomated safety check: NotesNone
Research Refine PipelinezjYao36/Auto-Research-Refine1285 repos~1.4kAutomated safety check: NotesNone
Metabolic Study Planneraiming-lab/AutoResearchClaw15k—~1.9kAutomated safety check: PassMIT

Similar skills

  • Scientific Critical Thinking

    weapp-tailwindcss/weapp-tailwindcss

    Evaluate research rigor. An agent skill from weapp-tailwindcss/weapp-tailwindcss.

    1.9k GitHub starsUsed in 22 repos~5.9k tokens
    Research & ScienceAuto-check: notes
  • Benchmark Paper Template

    HKUSTDial/Supervisor-Skills

    Structures benchmark and evaluation papers around five pillars, with a completeness audit, an Introduction logic chain, a section skeleton and a pre-submission checklist.

    8.7k GitHub stars~2.8k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Claim-Driven Experiment Planner

    zjYao36/Auto-Research-Refine

    Turns a refined research proposal into a claim-to-evidence-to-run-order roadmap instead of a sprawling benchmark wishlist.

    128 GitHub starsUsed in 6 repos~2.3k tokens
    Research & ScienceAuto-check: notes
  • Research Refine Pipeline

    zjYao36/Auto-Research-Refine

    Chains research-refine and experiment-plan to turn a vague research direction into a focused proposal and a claim-driven experiment roadmap.

    128 GitHub starsUsed in 5 repos~1.4k tokens
    Research & ScienceAuto-check: notes
  • Metabolic Study Planner

    aiming-lab/AutoResearchClaw

    Turns a broad metabolic modelling topic into a concrete, paper-shaped plan with organism, model, perturbations, metrics and figures before any FBA code is written.

    15k GitHub stars~1.9k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Analytical Method Validation Planner

    K-Dense-AI/scientific-agent-skills

    Plans, runs, and documents analytical method validation, verification, or transfer studies under ICH Q2(R2)/Q14, USP, ICH M10, CLSI EP, or ISO/IEC 17025.

    48k GitHub starsUsed in 1 repo~4.9k tokens
    Research & ScienceAuto-check: notes

More from flonat/flonat-research

All 83 skills in this repo
  • Latex Posters

    flonat/flonat-research

    Create a large-format academic poster in LaTeX using beamerposter, tikzposter, or baposter.

    146 GitHub stars~1.5k tokensUpdated 9 days ago
    Auto-check: notes
  • Skill Creator

    flonat/flonat-research

    Create, revise, and evaluate reusable AI workflow skills, including trigger-quality tests.

    146 GitHub stars~4.4k tokensUpdated 9 days ago
    Auto-check passed
  • DOCX

    flonat/flonat-research

    Create, read, edit, or convert Microsoft Word documents while preserving professional document structure.

    146 GitHub stars~1.2k tokensUpdated 9 days ago
    Auto-check passed
  • PDF

    flonat/flonat-research

    Read, create, combine, split, rotate, OCR, watermark, secure, or extract content from PDF files.

    146 GitHub stars~488 tokensUpdated 9 days ago
    Auto-check passed
  • Init Project Orchestration

    flonat/flonat-research

    Create or migrate project-level agents, repeatable project workflows, and planning state from one client-neutral contract, then render repository-scoped adapters for both Claude Code and Codex.

    146 GitHub stars~1.6k tokensUpdated 9 days ago
    Auto-check passed
  • Pre Commit Audit

    flonat/flonat-research

    Deliver a fast pre-commit safety scan: file size, anonymity (author / affiliation strings in tex/bib), hardcoded secrets, and invisible-Unicode carriers.

    146 GitHub stars~2.8k tokensUpdated 9 days ago
    Auto-check: notes

Questions about Experiment Design

What does Experiment Design do?

Design empirical studies through power analysis, pre-analysis planning, QSF parsing, and survey architecture. Experiment Design is an agent skill from flonat/flonat-research. Design empirical studies through power analysis, pre-analysis planning, QSF parsing, and survey architecture.

When should I use Experiment Design?

Experiment Design fits situations like: specifying sampling; analysis before data collection.

How do I install Experiment Design in Claude Code?

Run `npx skills add flonat/flonat-research --skill experiment-design -a claude-code`. Or copy the skill folder (skills/experiment-design in flonat/flonat-research) into .claude/skills/experiment-design in your project. Claude Code loads it when a task matches its description.

How do I install Experiment Design in Codex?

Run `npx skills add flonat/flonat-research --skill experiment-design -a codex`. Or copy the skill folder (skills/experiment-design in flonat/flonat-research) into .agents/skills/experiment-design in your project. Codex loads it when a task matches its description.

Can I use Experiment Design in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add flonat/flonat-research --skill experiment-design -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/experiment-design, .gemini/skills/experiment-design, .github/skills/experiment-design and .opencode/skills/experiment-design in your project.

What does Experiment Design need to run?

SKILL.md names no scripts, command-line tools or credentials: Experiment Design is instructions for the agent only. Our summary lists: Python 3. Its frontmatter pre-approves these tools: Bash(uv*, Rscript*, R*, mkdir*, ls*), Read, Write, Edit, Glob, Grep, AskUserQuestion.

Does Experiment Design access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Experiment Design safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Experiment Design use?

Experiment Design is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Experiment Design use?

About 2.3k tokens (SKILL.md is roughly 9.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 6.9k tokens, read only when the agent opens those files.

What are the alternatives to Experiment Design?

Skills that share tags, products or a category with Experiment Design: Scientific Critical Thinking (weapp-tailwindcss/weapp-tailwindcss, 1.9k stars), Benchmark Paper Template (HKUSTDial/Supervisor-Skills, 8.7k stars), Claim-Driven Experiment Planner (zjYao36/Auto-Research-Refine, 128 stars) and Research Refine Pipeline (zjYao36/Auto-Research-Refine, 128 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Experiment Design?

flonat (a GitHub user) maintains it in flonat/flonat-research, which has 146 GitHub stars. The repository holds 83 skills in this directory. The repository was last updated on September 29, 2026.

Source: flonat/flonat-research on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.