Agent skill

Statistical Problem Formulation

by aiming-lab in aiming-lab/AutoResearchClaw

Formulate statistical research problems with formal notation, target parameters, assumptions, hypotheses, evaluation criteria, and theory targets.

MITAuto-check passedResearch & Science

Install Statistical Problem Formulation

skills CLI
$ npx skills add aiming-lab/AutoResearchClaw --skill statistical-problem-formulation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install aiming-lab/AutoResearchClaw statistical-problem-formulation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/aiming-lab/AutoResearchClaw.git skills-src && mkdir -p .claude/skills && cp -r skills-src/external/agents/stat_research_agent/skills/statistical-problem-formulation .claude/skills/statistical-problem-formulation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
statistical-problem-formulation
GitHub stars
15k
Token cost
~671 tokens
SKILL.md length
162 words
Files
1
Skills in repo
34
Repo updated
First seen
Licence
MIT

At a glance

Formulate statistical research problems with formal notation, target parameters, assumptions, hypotheses, evaluation criteria, and theory targets.

  • Research & Science work in your project
  • SKILL.md covers Overview, Required Formulation Elements, Handoff Schema and Template, plus 1 more section
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Statistical Problem Formulation is an agent skill from aiming-lab/AutoResearchClaw. Formulate statistical research problems with formal notation, target parameters, assumptions, hypotheses, evaluation criteria, and theory targets.

Its SKILL.md is about 670 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Research & Science. The repository describes itself as: Fully autonomous & self-evolving research from idea to paper. Chat an Idea. Get a Paper. 🦞. The licence is MIT.

When your agent uses it

  • Research & Science work in your project

Example prompts

  • “/statistical-problem-formulation”

What it can do on your machine

Read from SKILL.md and the folder at commit be4ba47. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are yaml and markdown).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Statistical Problem Formulation loads about 671 tokens when it runs. Until then it costs about 45 tokens; SKILL.md has 162 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~45
When it runs · the whole SKILL.md, loaded when a task matches
~671

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from aiming-lab/AutoResearchClaw at commit be4ba47, republished under its MIT licence (© aiming-lab). 162 words, ~671 tokens.

Download SKILL.mdSave it as .claude/skills/statistical-problem-formulation/SKILL.md (or your agent's skills folder).
name
statistical-problem-formulation
description
Formulate statistical research problems with formal notation, target parameters, assumptions, hypotheses, evaluation criteria, and theory targets.
metadata.category
domain
metadata.trigger-keywords
problem formulation,statistical formulation,estimand,assumptions,data model,hypothesis,theory target
metadata.applicable-stages
1,2,3,4,5
metadata.priority
1

Statistical Problem Formulation

Overview

Use this skill before any method design, theory, experiment, or report writing. The goal is to transform a broad topic into a precise statistical problem.

Required Formulation Elements

ElementQuestions
Observed dataWhat is observed? What is the sample size? Are samples iid, dependent, clustered, censored, or selected?
Data modelWhat family of distributions or data-generating processes is considered?
TargetWhat parameter, decision, prediction, or risk is the object of study?
AssumptionsWhat must hold for the target to be identifiable or the method to work?
HypothesesWhat claims should be supported, refuted, or made inconclusive?
CriteriaWhat metrics define success or failure?
Theory targetWhat property should be derived: bias, variance, consistency, rate, coverage, error bound, robustness, or impossibility?

Handoff Schema

The problem formulation should be precise enough to support this structured handoff:

yaml
topic_id: TXX
title: ""
research_question: ""
observed_data:
  notation: ""
  sampling: iid | dependent | clustered | time_series | selected | unknown
data_model:
  notation: ""
  family: ""
target:
  name: ""
  notation: ""
  type: estimand | decision | prediction | risk | descriptive_quantity
  truth_source: analytic | simulation | oracle | empirical_reference | not_applicable
assumptions:
  structural: []
  sampling: []
  regularity: []
  identifiability: []
claims:
  - id: C1
    statement: ""
    formal_statement: ""
evaluation_criteria:
  - name: ""
    direction: ""
theory_targets:
  - identifiability
  - bias
  - consistency
blocking_ambiguities: []

Template

markdown
# Problem Formulation

## Research Question
...

## Observed Data
Let ...

## Data-Generating Model
Assume ...

## Target / Estimand
Define ...

## Candidate Procedure Class
We consider procedures ...

## Assumptions
1. ...

## Claims / Hypotheses
- ...

## Evaluation Criteria
- ...

## Theoretical Questions
- ...

## Experimental Questions
- ...

Quality Bar

A formulation passes only if another researcher could implement or analyze the problem without guessing the target, assumptions, or success criteria.

© aiming-lab, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in external/agents/stat_research_agent/skills/statistical-problem-formulation of aiming-lab/AutoResearchClaw.

Open the folder on GitHubat commit be4ba47

Compare with similar skills

Statistical Problem Formulation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Statistical Problem Formulation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Statistical Problem Formulation this skillaiming-lab/AutoResearchClaw15k—~671Automated safety check: PassMIT
Hypothesis Generationspacering-net/codeg3.8k15 repos~3.6kAutomated safety check: NotesMIT
GitHub Deep Researchbytedance/deer-flow83k5 repos~1.3kAutomated safety check: PassMIT
Nature Paper CardYuan1z0825/nature-skills46k2 repos~2.1kAutomated safety check: PassApache-2.0
Read arXiv Paperkarpathy/nanochat58k2 repos~494Automated safety check: PassMIT
Content Research Writerweapp-tailwindcss/weapp-tailwindcss1.9k25 repos~3.5kAutomated safety check: PassMIT

Similar skills

  • Hypothesis Generation

    spacering-net/codeg

    Structured hypothesis formulation from observations. An agent skill from spacering-net/codeg.

    3.8k GitHub starsUsed in 15 repos~3.6k tokens
    Research & ScienceAuto-check: notes
  • GitHub Deep Research

    bytedance/deer-flow

    Researches a GitHub repository over four rounds using the GitHub API and web search, then writes a structured markdown report with timeline, metrics and Mermaid diagrams.

    83k GitHub starsUsed in 5 repos~1.3k tokens
    Research & ScienceAuto-check passed
  • Nature Paper Card

    Yuan1z0825/nature-skills

    Builds a structured deep-reading card for one scientific paper, covering methods, how experiments support claims, limitations and research ideas, with a script to prepare the source.

    46k GitHub starsUsed in 2 repos~2.1k tokens
    Research & ScienceAuto-check passed
  • Read arXiv Paper

    karpathy/nanochat

    Fetches the TeX source of an arXiv paper from its URL, reads it and writes a markdown summary tied to the nanochat project.

    58k GitHub starsUsed in 2 repos~494 tokens
    Research & ScienceAuto-check passed
  • Content Research Writer

    weapp-tailwindcss/weapp-tailwindcss

    Assists in writing high-quality content by conducting research, adding citations, improving hooks, iterating on outlines, and providing real-time feedback on each section.

    1.9k GitHub starsUsed in 25 repos~3.5k tokens
    Research & ScienceAuto-check passed
  • Peer Review

    spacering-net/codeg

    Structured manuscript/grant review with checklist-based evaluation.

    3.8k GitHub starsUsed in 18 repos~5.9k tokens
    Research & ScienceAuto-check: notes

More from aiming-lab/AutoResearchClaw

All 34 skills in this repo
  • A-Evolve Agent Improvement

    aiming-lab/AutoResearchClaw

    Diagnoses where an agent failed across runs and turns the findings into new skills, system prompt patches and knowledge entries, using the A-Evolve loop.

    15k GitHub stars~1.8k tokensUpdated 1 mo ago
    Auto-check passed
  • Metabolic Study Planner

    aiming-lab/AutoResearchClaw

    Turns a broad metabolic modelling topic into a concrete, paper-shaped plan with organism, model, perturbations, metrics and figures before any FBA code is written.

    15k GitHub stars~1.9k tokensUpdated 1 mo ago
    Auto-check passed
  • MFA Pipeline Orchestrator

    aiming-lab/AutoResearchClaw

    Runs a metabolic flux analysis from model loading to phenotype prediction and figures by handing work to four sub-agents in sequence.

    15k GitHub stars~923 tokensUpdated 1 mo ago
    Auto-check passed
  • Qiskit 2.x Quantum ML Reference

    aiming-lab/AutoResearchClaw

    Reference patterns for writing qiskit 2.x code for variational quantum machine learning: feature maps, VQC training, VQE for chemistry, MPS circuits and noise models.

    15k GitHub stars~4.7k tokensUpdated 1 mo ago
    Auto-check passed
  • Genome-Scale Metabolic Model Builder

    aiming-lab/AutoResearchClaw

    Builds or loads a genome-scale metabolic model in COBRApy, sets its growth medium and objective, and exports it as a validated JSON file for flux analysis.

    15k GitHub stars~1.4k tokensUpdated 1 mo ago
    Auto-check passed
  • Biopython Bioinformatics

    aiming-lab/AutoResearchClaw

    Quick reference for Biopython work: sequence operations, SeqIO file parsing, BLAST searches, Entrez queries, phylogenetic trees and PDB structure analysis.

    15k GitHub stars~810 tokensUpdated 1 mo ago
    Auto-check passed

Questions about Statistical Problem Formulation

What does Statistical Problem Formulation do?

Formulate statistical research problems with formal notation, target parameters, assumptions, hypotheses, evaluation criteria, and theory targets. Statistical Problem Formulation is an agent skill from aiming-lab/AutoResearchClaw. Formulate statistical research problems with formal notation, target parameters, assumptions, hypotheses, evaluation criteria, and theory targets.

When should I use Statistical Problem Formulation?

Statistical Problem Formulation fits situations like: research & Science work in your project.

How do I install Statistical Problem Formulation in Claude Code?

Run `npx skills add aiming-lab/AutoResearchClaw --skill statistical-problem-formulation -a claude-code`. Or copy the skill folder (external/agents/stat_research_agent/skills/statistical-problem-formulation in aiming-lab/AutoResearchClaw) into .claude/skills/statistical-problem-formulation in your project. Claude Code loads it when a task matches its description.

How do I install Statistical Problem Formulation in Codex?

Run `npx skills add aiming-lab/AutoResearchClaw --skill statistical-problem-formulation -a codex`. Or copy the skill folder (external/agents/stat_research_agent/skills/statistical-problem-formulation in aiming-lab/AutoResearchClaw) into .agents/skills/statistical-problem-formulation in your project. Codex loads it when a task matches its description.

Can I use Statistical Problem Formulation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add aiming-lab/AutoResearchClaw --skill statistical-problem-formulation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/statistical-problem-formulation, .gemini/skills/statistical-problem-formulation, .github/skills/statistical-problem-formulation and .opencode/skills/statistical-problem-formulation in your project.

What does Statistical Problem Formulation need to run?

SKILL.md names no scripts, command-line tools or credentials: Statistical Problem Formulation is instructions for the agent only.

Does Statistical Problem Formulation access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Statistical Problem Formulation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Statistical Problem Formulation use?

Statistical Problem Formulation is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Statistical Problem Formulation use?

About 671 tokens (SKILL.md is roughly 2.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Statistical Problem Formulation?

Skills that share tags, products or a category with Statistical Problem Formulation: Hypothesis Generation (spacering-net/codeg, 3.8k stars), GitHub Deep Research (bytedance/deer-flow, 83k stars), Nature Paper Card (Yuan1z0825/nature-skills, 46k stars) and Read arXiv Paper (karpathy/nanochat, 58k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Statistical Problem Formulation?

aiming-lab (a GitHub organization) maintains it in aiming-lab/AutoResearchClaw, which has 14,587 GitHub stars. The repository holds 34 skills in this directory. The repository was last updated on August 19, 2026.

Source: aiming-lab/AutoResearchClaw on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.