Agent skill

Experiment Results Analysis for Papers

by LigphiDonk in LigphiDonk/Oh-my--paper

Turns experimental data such as CSV, JSON or TensorBoard logs into statistical significance tests, visualizations and a drafted Results section.

MITAuto-check passedData & Analytics

Install Experiment Results Analysis for Papers

skills CLI
$ npx skills add LigphiDonk/Oh-my--paper --skill inno-experiment-analysis -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install LigphiDonk/Oh-my--paper inno-experiment-analysis --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/LigphiDonk/Oh-my--paper.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/inno-experiment-analysis .claude/skills/inno-experiment-analysis && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
inno-experiment-analysis
GitHub stars
738
Token cost
~3k tokens
SKILL.md length
1,054 words
Files
8 (incl. references)
Skills in repo
27
Repo updated
First seen
Licence
MIT

At a glance

Turns experimental data such as CSV, JSON or TensorBoard logs into statistical significance tests, visualizations and a drafted Results section.

  • Works in 6 steps: Data Loading and Validation → Statistical Analysis → Model Performance Comparison → …
  • Analyzing experimental results across multiple model runs
  • SKILL.md covers Canonical Summary, Trigger Rules, Resource Use Rules and Execution Contract, plus 8 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

The pipeline runs data loading and validation, statistical analysis, visualization, writing and a final quality check, in that order. Supported input formats include CSV and JSON files, TensorBoard training-curve logs and Python pickle objects, validated for completeness (missing values, outliers), consistency (format and units) and reproducibility (recorded random seeds and version info) before any analysis runs.

It performs statistical significance tests and compares performance across multiple models, builds publication-quality visualizations from the validated data, and generates text for a paper's Results section grounded in that analysis. Bundled reference files cover statistical methods, results-writing guidance, visualization best practices and common pitfalls, with worked examples showing a finished analysis report and a finished results section side by side; generated outputs are saved into the active project rather than back into the skill's own directory.

When your agent uses it

  • Analyzing experimental results across multiple model runs
  • Generating the Results section of a research paper from data
  • Running statistical significance tests to compare model performance

Example prompts

  • “Analyze these three models' benchmark CSVs and test for significance.”
  • “Write the Results section from this experiment's TensorBoard logs.”
  • “Check these experimental results for missing values or reproducibility issues.”

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Data Loading and Validation
  2. Statistical Analysis
  3. Model Performance Comparison
  4. Visualization
  5. Writing the Results Section
  6. Quality Check

What it can do on your machine

Read from SKILL.md and the folder at commit 6baece9. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are markdown).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • nature.com
    • science.org
    • neurips.cc

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Experiment Results Analysis for Papers loads about 3k tokens when it runs, and up to ~14k if it reads all its reference files. Until then it costs about 67 tokens; SKILL.md has 1,054 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~67
When it runs · the whole SKILL.md, loaded when a task matches
~3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~14k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from LigphiDonk/Oh-my--paper at commit 6baece9, republished under its MIT licence (© LigphiDonk). 1,054 words, ~2,989 tokens.

Download SKILL.mdSave it as .claude/skills/inno-experiment-analysis/SKILL.md (or your agent's skills folder). This skill also uses 7 other files; get the full folder from GitHub.
name
inno-experiment-analysis
description
This skill should be used when the user asks to "analyze experimental results", "generate results section", "statistical analysis of experiments", "compare model performance", "create results visualization", or mentions connecting experime...
id
inno-experiment-analysis
version
0.1.0
stages
experiment
tools
read_file, search_project, write_file
summary
This skill should be used when the user asks to "analyze experimental results", "generate results section", "statistical analysis of experiments", "compare…
primaryIntent
evaluation
intents
evaluation, experiment
capabilities
evaluation-benchmarking
domains
general
keywords
research, analysis, statistics, visualization, paper writing, inno-experiment-analysis, evaluation-benchmarking, inno, experiment, this, should, be

inno-experiment-analysis

Canonical Summary

This skill should be used when the user asks to "analyze experimental results", "generate results section", "statistical analysis of experiments", "compare model performance", "create results visualization", or mentions connecting experime...

Trigger Rules

Use this skill when the user request matches its research workflow scope. Prefer the bundled resources instead of recreating templates or reference material. Keep outputs traceable to project files, citations, scripts, or upstream evidence.

Resource Use Rules

  • Read from references/ only when the current task needs the extra detail.

Execution Contract

  • Resolve every relative path from this skill directory first.
  • Prefer inspection before mutation when invoking bundled scripts.
  • If a required runtime, CLI, credential, or API is unavailable, explain the blocker and continue with the best manual fallback instead of silently skipping the step.
  • Do not write generated artifacts back into the skill directory; save them inside the active project workspace.

Upstream Instructions

Results Analysis for ML/AI Research

A systematic experimental results analysis workflow connecting experimental data to paper writing.

Core Features

This skill provides three core capabilities:

  1. Experimental Data Analysis - Read and analyze experimental data in various formats
  2. Statistical Validation - Perform statistical significance tests and performance comparisons
  3. Paper Content Generation - Generate text and visualizations for the Results section

When to Use

Use this skill when you need to:

  • Analyze experimental results (CSV, JSON, TensorBoard logs)
  • Generate the Results section of a paper
  • Compare performance across multiple models
  • Perform statistical significance tests
  • Create publication-quality visualizations
  • Validate the reliability of experimental results

Workflow

Standard Analysis Pipeline
Data Loading → Data Validation → Statistical Analysis → Visualization → Writing → Quality Check
Step 1: Data Loading and Validation

Supported Data Formats:

  • CSV files - Tabular data
  • JSON files - Structured results
  • TensorBoard logs - Training curves
  • Python pickle - Complex objects

Data Validation Checks:

  • Completeness check - Missing values, outliers
  • Consistency check - Data format, units
  • Reproducibility check - Random seeds, version info

Select appropriate tools for data loading and preliminary validation based on data format.

Step 2: Statistical Analysis

Basic Statistics:

  • Mean
  • Standard Deviation
  • Standard Error
  • Confidence Interval

Significance Tests:

  • t-test - Two-group comparison
  • ANOVA - Multi-group comparison
  • Wilcoxon test - Non-parametric test
  • Bonferroni correction - Multiple comparison correction

Select appropriate statistical tests based on data characteristics.

Key Principles:

  • Report complete statistical information (mean ± std/SE)
  • Specify the test method and significance level used
  • Report p-values and effect sizes
  • Consider multiple comparison issues

See references/statistical-methods.md for the complete statistical methods guide.

Step 3: Model Performance Comparison

Comparison Dimensions:

  • Accuracy/Performance metrics
  • Training time/Inference speed
  • Model complexity/Parameter count
  • Robustness/Generalization ability

Comparison Methods:

  • Baseline comparison - Compare with existing methods
  • Ablation study - Validate component contributions
  • Cross-dataset validation - Test generalization

Systematically compare performance across different methods, ensuring fair comparison.

Step 4: Visualization

Publication-Quality Visualization Requirements:

  • Vector format (PDF/EPS)
  • Colorblind-friendly palette
  • Clear labels and legends
  • Appropriate error bars
  • Readable in black-and-white print

Common Chart Types:

  • Line chart - Training curves, trend analysis
  • Bar chart - Performance comparison
  • Box plot - Distribution display
  • Heatmap - Correlation analysis
  • Scatter plot - Relationship display

Use appropriate visualization tools to generate publication-quality figures.

See references/visualization-best-practices.md for the visualization guide.

Step 5: Writing the Results Section

Results Section Structure:

markdown
## Results

### Overview of Main Findings
[1-2 paragraphs summarizing core results]

### Experimental Setup
[Brief description of experimental configuration; details in appendix]

### Performance Comparison
[Comparison with baseline methods, including tables and figures]

### Ablation Study
[Validate contributions of each component]

### Statistical Significance
[Report statistical test results]

### Qualitative Analysis
[Case studies, visualization examples]

Writing Principles:

  • Clearly state the hypothesis each experiment validates
  • Guide readers to observe key phenomena: "Figure X shows..."
  • Report complete statistical information
  • Honestly report limitations

See references/results-writing-guide.md for the complete writing guide.

Step 6: Quality Check

Checklist:

  • All values include error bars/confidence intervals
  • Statistical test methods are specified
  • Figures are clear and readable (including black-and-white print)
  • Hyperparameter search ranges are reported
  • Computational resources are specified (GPU type, time)
  • Random seed settings are specified
  • Results are reproducible (code/data available)

Common Mistakes and Pitfalls

Statistical Errors

❌ Wrong approach:

  • Reporting only the best results (cherry-picking)
  • Confusing standard deviation and standard error
  • Not reporting statistical significance
  • Not correcting for multiple comparisons

✅ Correct approach:

  • Report all experimental results
  • Clearly specify whether standard deviation or standard error is used
  • Perform appropriate statistical tests
  • Use Bonferroni or similar correction methods
Show full SKILL.md (428 more words)Show less
Visualization Errors

❌ Wrong approach:

  • Using non-colorblind-friendly palettes
  • Y-axis not starting from 0 (exaggerating differences)
  • Missing error bars
  • Overly complex figures

✅ Correct approach:

  • Use Okabe-Ito or Paul Tol palettes
  • Set reasonable axis ranges
  • Include error bars and confidence intervals
  • Keep figures clean and clear
Writing Errors

❌ Wrong approach:

  • Over-interpreting results
  • Not describing experimental setup
  • Hiding negative results
  • Missing statistical information

✅ Correct approach:

  • Objectively describe observed phenomena
  • Provide sufficient experimental details
  • Honestly report all results
  • Report complete statistical information

See references/common-pitfalls.md for the complete error patterns and fixes.

Integration with Paper Writing

Collaboration with ml-paper-writing Skill

This skill focuses on experimental results analysis and works in tandem with the ml-paper-writing skill:

inno-experiment-analysis handles:

  • Data analysis and statistical tests
  • Visualization generation
  • Results interpretation

ml-paper-writing handles:

  • Complete paper structure
  • Citation management
  • Conference format requirements

Workflow Integration:

Experiments complete → inno-experiment-analysis analyzes
    ↓
Generate analysis report and visualizations
    ↓
ml-paper-writing integrates into paper
    ↓
Complete Results section
Output Format

After analysis, the following are generated:

  1. Analysis Report (analysis-report.md)

    • Statistical summary
    • Key findings
    • Suggested figures
  2. Visualization Files (figures/)

    • PDF format figures
    • Standalone figure captions
  3. Results Draft (results-draft.md)

    • Text ready for direct use in the paper
    • Includes figure references

Examples and Templates

Example Files

Refer to the examples/ directory for complete examples:

  • example-analysis-report.md - Complete analysis report example
  • example-results-section.md - Paper Results section example
Workflow Overview

The complete analysis pipeline includes:

  1. Data Loading - Read results from experiment output files
  2. Statistical Analysis - Compute basic statistics and perform significance tests
  3. Visualization - Create publication-quality figures
  4. Report Generation - Integrate analysis results and visualizations

See the guides in the references/ directory for detailed methods and best practices.

Reference Resources

Detailed Guides
  • references/statistical-methods.md - Complete statistical methods guide
  • references/results-writing-guide.md - Results section writing standards
  • references/visualization-best-practices.md - Visualization best practices
  • references/common-pitfalls.md - Common errors and fixes
External Resources

Best Practices Summary

Data Analysis

✅ Recommended:

  • Run experiments multiple times (at least 3-5 runs)
  • Report complete statistical information
  • Use appropriate statistical tests
  • Check data completeness

❌ Prohibited:

  • Cherry-picking best results
  • Ignoring statistical significance
  • Hiding negative results
  • Not reporting experimental setup
Visualization

✅ Recommended:

  • Use vector format
  • Colorblind-friendly palettes
  • Include error bars
  • Clear labels

❌ Prohibited:

  • Raster formats (PNG/JPG)
  • Misleading axis scales
  • Overly complex figures
  • Missing legends
Writing

✅ Recommended:

  • Objectively describe results
  • Provide sufficient detail
  • Honestly report limitations
  • Guide reader attention

❌ Prohibited:

  • Over-interpretation
  • Hiding details
  • Exaggerating effects
  • Vague descriptions

Summary

This skill provides a systematic experimental results analysis workflow:

  1. Data Loading and Validation - Ensure data quality
  2. Statistical Analysis - Perform appropriate statistical tests
  3. Model Comparison - Systematic performance comparison
  4. Visualization - Publication-quality figures
  5. Writing - Results section content
  6. Quality Check - Ensure reproducibility

Following these principles produces high-quality, reproducible experimental results analysis that meets top conference standards.

© LigphiDonk, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 7 other files (references) in skills/inno-experiment-analysis of LigphiDonk/Oh-my--paper.

  • SKILL.md
  • USAGE.md
  • examples/example-analysis-report.md
  • examples/example-results-section.md
  • references/common-pitfalls.md
  • references/results-writing-guide.md
  • references/statistical-methods.md
  • references/visualization-best-practices.md

Open the folder on GitHubat commit 6baece9

Compare with similar skills

Experiment Results Analysis for Papers next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Experiment Results Analysis for Papers compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Experiment Results Analysis for Papers this skillLigphiDonk/Oh-my--paper738—~3kAutomated safety check: PassMIT
Analysis Graphingclshortfuse/renodx4.5k—~1.1kAutomated safety check: PassMIT
CSV Data Analysis5zjk5/prompt-engineering127—~2.6kAutomated safety check: PassNone
Data Analysisfastclaw-ai/fastclaw1.4k—~410Automated safety check: PassCustom licence
Results AnalysisGalaxy-Dawn/claude-scholar5.7k1 repos~2.4kAutomated safety check: PassMIT
Data Analystholaboss-ai/holaOS11k—~551Automated safety check: PassCustom licence

Similar skills

  • Analysis Graphing

    clshortfuse/renodx

    RenoDX workflow for creating readable analysis graphs and plots from shader math, CSVs, EXRs, LUTs, hue sweeps, tone curves, gamut comparisons, energy/scalar maps, and test-pattern statistics.

    4.5k GitHub stars~1.1k tokensUpdated yesterday
    Data & AnalyticsAuto-check passed
  • CSV Data Analysis

    5zjk5/prompt-engineering

    This skill should be used when users need to analyze CSV or Excel files, understand data patterns, generate statistical summaries, or create data visualizations.

    127 GitHub stars~2.6k tokensUpdated 24 days ago
    Data & AnalyticsAuto-check passed
  • Data Analysis

    fastclaw-ai/fastclaw

    Analyze data, process CSV/JSON files, compute statistics, and create data visualizations.

    1.4k GitHub stars~410 tokensUpdated 3 days ago
    Data & AnalyticsAuto-check passed
  • Results Analysis

    Galaxy-Dawn/claude-scholar

    This skill should be used when the user asks to "analyze experimental results", "run strict statistical analysis", "compare model performance", "generate scientific figures", "check significance"…

    5.7k GitHub starsUsed in 1 repo~2.4k tokens
    Data & AnalyticsAuto-check passed
  • Data Analyst

    holaboss-ai/holaOS

    Analyzes a dataset, spreadsheet or metrics table, reports what changed and what is driving it, and recommends which chart to use for each key finding.

    11k GitHub stars~551 tokensUpdated 1 mo ago
    Data & AnalyticsAuto-check passed
  • Scientific Toolkit Skill

    zLanqing/codex-claude-academic-skills

    Research computing toolkit for optoelectronic information science and engineering, MATLAB/Octave, Python scientific analysis, signal processing, image processing, statistics, simulation…

    4.7k GitHub stars~1.2k tokensUpdated 4 mo ago
    Data & AnalyticsAuto-check passed

More from LigphiDonk/Oh-my--paper

All 27 skills in this repo
  • Preprint Search on bioRxiv

    LigphiDonk/Oh-my--paper

    Searches bioRxiv life sciences preprints by keyword, author, date range or category with a Python script, returning JSON metadata and optional PDF downloads.

    738 GitHub starsUsed in 12 repos~3.7k tokens
    Auto-check passed
  • Literature PDF OCR Library Builder

    LigphiDonk/Oh-my--paper

    Searches and downloads legally accessible academic PDFs, OCRs them to Markdown, and organizes the results into a traceable, AI-readable literature library.

    738 GitHub stars~1.1k tokensUpdated 5 mo ago
    Auto-check passed
  • Inno Code Survey

    LigphiDonk/Oh-my--paper

    Finds and clones missing code repositories for a chosen research idea, then writes a survey that maps academic concepts to their implementations.

    738 GitHub stars~3.6k tokensUpdated 5 mo ago
    Auto-check passed
  • Citation Verification Guide

    LigphiDonk/Oh-my--paper

    Lays out principles for catching fake, mismatched, or inconsistently formatted citations in academic writing, checked through live web search.

    738 GitHub stars~2.2k tokensUpdated 5 mo ago
    Auto-check passed
  • Single-Cell Initial Analysis

    LigphiDonk/Oh-my--paper

    Runs a seven-step quality-control and exploration pipeline on scRNA-seq, CyTOF or flow cytometry data and writes a plain-language report of what it found.

    738 GitHub starsUsed in 1 repo~1.4k tokens
    Auto-check passed
  • Making Academic Presentations

    LigphiDonk/Oh-my--paper

    Create academic presentation slide decks and optionally demo videos from research papers.

    738 GitHub stars~2k tokensUpdated 5 mo ago
    Auto-check passed

Questions about Experiment Results Analysis for Papers

What does Experiment Results Analysis for Papers do?

Turns experimental data such as CSV, JSON or TensorBoard logs into statistical significance tests, visualizations and a drafted Results section. The pipeline runs data loading and validation, statistical analysis, visualization, writing and a final quality check, in that order. Supported input formats include CSV and JSON files, TensorBoard training-curve logs and Python pickle objects, validated for completeness (missing values, outliers), consistency (format and units) and reproducibility (recorded random seeds and version info) before any analysis runs.

When should I use Experiment Results Analysis for Papers?

Experiment Results Analysis for Papers fits situations like: analyzing experimental results across multiple model runs; generating the Results section of a research paper from data; running statistical significance tests to compare model performance.

How do I install Experiment Results Analysis for Papers in Claude Code?

Run `npx skills add LigphiDonk/Oh-my--paper --skill inno-experiment-analysis -a claude-code`. Or copy the skill folder (skills/inno-experiment-analysis in LigphiDonk/Oh-my--paper) into .claude/skills/inno-experiment-analysis in your project. Claude Code loads it when a task matches its description.

How do I install Experiment Results Analysis for Papers in Codex?

Run `npx skills add LigphiDonk/Oh-my--paper --skill inno-experiment-analysis -a codex`. Or copy the skill folder (skills/inno-experiment-analysis in LigphiDonk/Oh-my--paper) into .agents/skills/inno-experiment-analysis in your project. Codex loads it when a task matches its description.

Can I use Experiment Results Analysis for Papers in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add LigphiDonk/Oh-my--paper --skill inno-experiment-analysis -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/inno-experiment-analysis, .gemini/skills/inno-experiment-analysis, .github/skills/inno-experiment-analysis and .opencode/skills/inno-experiment-analysis in your project.

What does Experiment Results Analysis for Papers need to run?

SKILL.md names no scripts, command-line tools or credentials: Experiment Results Analysis for Papers is instructions for the agent only.

Does Experiment Results Analysis for Papers access the network?

SKILL.md names 3 domains. As links in the text: nature.com, science.org and neurips.cc. This is read from the text; nothing was executed.

Is Experiment Results Analysis for Papers safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Experiment Results Analysis for Papers use?

Experiment Results Analysis for Papers is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Experiment Results Analysis for Papers use?

About 3k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 11k tokens, read only when the agent opens those files.

What are the alternatives to Experiment Results Analysis for Papers?

Skills that share tags, products or a category with Experiment Results Analysis for Papers: Analysis Graphing (clshortfuse/renodx, 4.5k stars), CSV Data Analysis (5zjk5/prompt-engineering, 127 stars), Data Analysis (fastclaw-ai/fastclaw, 1.4k stars) and Results Analysis (Galaxy-Dawn/claude-scholar, 5.7k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Experiment Results Analysis for Papers?

LigphiDonk (a GitHub user) maintains it in LigphiDonk/Oh-my--paper, which has 738 GitHub stars. The repository holds 27 skills in this directory. The repository was last updated on April 15, 2026.

Source: LigphiDonk/Oh-my--paper on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.