Agent skill

Results Analysis

by Galaxy-Dawn in Galaxy-Dawn/claude-scholar

This skill should be used when the user asks to "analyze experimental results", "run strict statistical analysis", "compare model performance", "generate scientific figures", "check significance"…

MITAuto-check passedData & Analytics

Install Results Analysis

skills CLI
$ npx skills add Galaxy-Dawn/claude-scholar --skill results-analysis -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Galaxy-Dawn/claude-scholar results-analysis --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Galaxy-Dawn/claude-scholar.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/results-analysis .claude/skills/results-analysis && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
results-analysis
GitHub stars
5.7k
Used in
1 other repo
Token cost
~2.4k tokens
SKILL.md length
1,033 words
Files
11 (incl. references)
Skills in repo
34
Repo updated
First seen
Licence
MIT

At a glance

This skill should be used when the user asks to "analyze experimental results", "run strict statistical analysis", "compare model performance", "generate scientific figures", "check significance"…

  • Works in 6 steps: Inventory and validate artifacts → Lock the comparison questions → Run strict statistics → …
  • Asks to analyze experimental results
  • SKILL.md covers Core contract, Non-negotiable quality bar, Standard workflow and Output structure, plus 5 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Results Analysis is an agent skill from Galaxy-Dawn/claude-scholar. This skill should be used when the user asks to "analyze experimental results", "run strict statistical analysis", "compare model performance", "generate scientific figures", "check significance", "do ablation analysis", or mentions interpreting experiment data with rigorous statistics and visualization. It focuses on strict analysis bundles, not Results-section prose.

Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 12 other files, including reference files (for example `USAGE.md`, `examples/example-analysis-report.md` and `examples/example-figure-catalog.md`).

It sits in Data & Analytics, covering Statistics and Data visualization. The repository describes itself as: Semi-automated research assistant for academic research and software development. Supports Claude Code, Codex CLI, Kimi Code CLI, and OpenCode across ideation, coding… The licence is MIT.

When your agent uses it

  • Asks to analyze experimental results
  • Run strict statistical analysis
  • Compare model performance
  • Generate scientific figures

Example prompts

  • “analyze experimental results”
  • “run strict statistical analysis”
  • “compare model performance”
  • “/results-analysis”

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Inventory and validate artifacts
  2. Lock the comparison questions
  3. Run strict statistics
  4. Generate real scientific figures
  5. Write analysis artifacts
  6. Final QA gate

What it can do on your machine

Read from SKILL.md and the folder at commit 9037873. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are markdown).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Results Analysis loads about 2.4k tokens when it runs, and up to ~9.1k if it reads all its reference files. Until then it costs about 97 tokens; SKILL.md has 1,033 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~97
When it runs · the whole SKILL.md, loaded when a task matches
~2.4k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~9.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Galaxy-Dawn/claude-scholar at commit 9037873, republished under its MIT licence (© Galaxy-Dawn). 1,033 words, ~2,363 tokens.

Download SKILL.mdSave it as .claude/skills/results-analysis/SKILL.md (or your agent's skills folder). This skill also uses 10 other files; get the full folder from GitHub.
name
results-analysis
description
This skill should be used when the user asks to "analyze experimental results", "run strict statistical analysis", "compare model performance", "generate scientific figures", "check significance", "do ablation analysis", or mentions interpreting experiment data with rigorous statistics and visualization. It focuses on strict analysis bundles, not Results-section prose.
tags
Research, Analysis, Statistics, Visualization, Scientific Reporting
version
0.2.0

Results Analysis

Run strict, evidence-first experimental analysis for ML/AI research.

Use this skill to produce a strict analysis bundle:

  • analysis-report.md
  • stats-appendix.md
  • figure-catalog.md
  • figures/

When the user asks for review, audit, no-write, dry-run, or when inputs are incomplete, use read-only audit mode instead of producing files or figures. In that mode, output only valid/invalid statistics, blockers, claim candidates, and what evidence is missing. If invoked by /analyze-results, the command layer may write a blocker summary, but this skill should not create figures, reports, or polished conclusions from incomplete evidence.

Do not use this skill to draft a paper Results section or a full experiment wrap-up report. Those belong to ml-paper-writing or results-report.

Core contract

This skill is responsible for
  • validating experiment artifacts and comparison units,
  • running rigorous descriptive and inferential statistics,
  • generating real scientific figures when data/logs are available,
  • writing figure purposes, caption requirements, and interpretation checklists,
  • surfacing limits, blockers, and missing evidence explicitly.
This skill is not responsible for
  • paper-ready Results prose,
  • manuscript narrative polishing,
  • paper-ready figure/table packaging with pubfig / pubtab,
  • project-level experiment retrospectives.

If the user wants the complete post-experiment summary report, hand off to results-report after this bundle is ready. If the user wants publication-grade figures/tables, export parameters, publication QA, or figure/table redesign, hand off to publication-chart-skill.

Non-negotiable quality bar

  1. Prefer real figures over figure specs. If the data can be read, generate real figures. Do not stop at “recommended visualization”. Exception: in read-only audit mode, do not generate figures; describe what figure would be valid after evidence is complete.
  2. Never fabricate statistics. If sample size, seeds, or raw metrics are missing, state the blocker clearly.
  3. Report complete statistics. Do not report only best scores or only p-values.
  4. Interpret every main figure. Every major figure must have purpose, caption requirements, and post-figure interpretation notes.
  5. Separate evidence from prose. This skill produces analysis artifacts; it does not write manuscript sections.

Standard workflow

1. Inventory and validate artifacts

Start by identifying:

  • metric tables (csv, json, tsv, logs),
  • training curves and checkpoints,
  • seeds / repeated runs,
  • baselines, ablations, and comparison families,
  • evaluation protocol metadata.

Validate:

  • metric direction (higher/lower is better),
  • unit of analysis (run, subject, fold, dataset, seed),
  • number of runs / seeds,
  • missing values or silent failures,
  • comparability across methods.

If the comparison is not statistically valid, say so before continuing. Do not treat repeated subject × task rows, folds, windows, trials, or seeds as independent units unless the design justifies it. Common blocker: a subject × task summary table is usually a repeated-measure summary, not an independent subject-level sample. If subjects have multiple task rows or missing task cells, state that before any significance or winner claim.

2. Lock the comparison questions

Before running statistics, define the exact comparison questions:

  • Which method is compared to which baseline?
  • What is the primary metric?
  • What is the repeated-measure unit?
  • Which ablation or robustness questions matter?
  • Which findings are decision-changing?

Do not mix unrelated comparisons into one undifferentiated table.

3. Run strict statistics

Always produce:

  • descriptive statistics: mean ± std when appropriate,
  • 95% CI or another clearly justified interval,
  • run/seed counts,
  • significance tests with assumptions stated,
  • effect sizes,
  • multiple-comparison handling when several contrasts are reported.

Default expectation:

  • check parametric assumptions first,
  • use non-parametric fallback when assumptions fail,
  • state exactly what was tested and on what samples.

See:

  • references/statistical-methods.md
  • references/statistical-reporting.md
4. Generate real scientific figures

Produce actual figures whenever artifacts are available.

Minimum expectation for a non-trivial analysis bundle:

  • one main comparison figure,
  • one supporting figure (training dynamics / ablation / breakdown / error analysis),
  • one exact numeric summary table in markdown.

Every main figure must define:

  • figure purpose,
  • plotted variables,
  • error bar meaning,
  • caption requirements,
  • interpretation checklist.

See:

  • references/visualization-best-practices.md
  • references/figure-interpretation.md
5. Write analysis artifacts
Show full SKILL.md (424 more words)Show less
analysis-report.md

Summarize:

  • the analysis question,
  • key findings,
  • strongest supported comparisons,
  • main caveats,
  • what changed in the experimental understanding,
  • claim candidates that may later be used in reports or manuscript writing.

Each claim candidate should use this shape:

md
## Claim Candidates

- Claim:
  - Source evidence:
  - Allowed wording:
  - Forbidden stronger wording:
  - Uncertainty:
  - Next check:
  - Decision: keep | weaken | revise | discard
stats-appendix.md

Record:

  • descriptive statistics,
  • test choices,
  • assumptions checked,
  • effect sizes,
  • confidence intervals,
  • multiple comparison corrections,
  • explicit blockers and limitations.
figure-catalog.md

For each figure, record:

  • filename,
  • purpose,
  • data source,
  • caption draft requirements,
  • key observation,
  • interpretation checklist,
  • known caveats.
6. Final QA gate

Do not finish until all are true:

  • the primary comparison question is explicit,
  • sample size / seed count is stated,
  • inferential tests are justified,
  • effect sizes are reported for major contrasts,
  • real figures exist when data exists,
  • each figure has an interpretation note,
  • limitations and blockers are explicit,
  • each supported or strong claim candidate has evidence, uncertainty, and allowed wording,
  • over-strong manuscript wording is explicitly blocked when evidence is insufficient,
  • no manuscript-style Results draft is included.

Output structure

text
analysis-output/
├── analysis-report.md
├── stats-appendix.md
├── figure-catalog.md
└── figures/
    ├── figure-01-main-comparison.pdf
    ├── figure-02-ablation.pdf
    └── ...

Figure interpretation rule

For every major figure, answer all three questions:

  1. Why does this figure exist?
  2. What exactly should the reader notice?
  3. What does that observation change in our belief or next decision?

If a figure cannot answer question 3, it is probably decorative rather than scientific.

Read-only audit mode

Use this mode when:

  • the user asks to audit or review existing artifacts,
  • the environment is read-only,
  • the user forbids file writes or figure generation,
  • core evidence is missing.

Return:

  • analysis questions,
  • valid statistics,
  • invalid or unsafe statistics,
  • claim candidates with allowed and forbidden wording,
  • blockers before report/figure generation.

Do not create analysis-output/, figures, or reports in this mode. Quarantine any statistics file whose interpretation contradicts its own p-value, test method, unit of analysis, or comparison family. Do not reuse that file for claim wording until provenance is checked.

Failure mode policy

When inputs are incomplete, say so explicitly.

Examples:

  • no seed-level data -> descriptive summary only; inferential claims blocked,
  • no comparable baseline outputs -> no significance claim,
  • no readable logs -> cannot generate dynamics figure,
  • too few runs -> effect size may be unstable; report this limitation.
  • unclear unit of analysis -> no winner claim or significance claim,
  • analysis file with contradictory interpretation -> quarantine it until provenance is checked.

Never replace missing evidence with confident prose.

Reference files

Load only what is needed:

  • references/statistical-methods.md - test selection and assumptions
  • references/statistical-reporting.md - minimum reporting standard
  • references/visualization-best-practices.md - publication-quality figure rules
  • references/figure-interpretation.md - how to explain figures with evidence
  • references/analysis-depth.md - move from observation to mechanism and decision
  • references/common-pitfalls.md - common analysis and reporting failures
  • ../research-ideation/references/research-contract.md - shared claim candidate and claim strength contract

Example files

  • examples/example-analysis-report.md
  • examples/example-stats-appendix.md
  • examples/example-figure-catalog.md

© Galaxy-Dawn, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 10 other files (references) in skills/results-analysis of Galaxy-Dawn/claude-scholar.

  • SKILL.md
  • USAGE.md
  • examples/example-analysis-report.md
  • examples/example-figure-catalog.md
  • examples/example-stats-appendix.md
  • references/analysis-depth.md
  • references/common-pitfalls.md
  • references/figure-interpretation.md
  • references/statistical-methods.md
  • references/statistical-reporting.md
  • references/visualization-best-practices.md

Open the folder on GitHubat commit 9037873

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in Galaxy-Dawn/claude-scholar, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Results Analysis next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Results Analysis compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Results Analysis this skillGalaxy-Dawn/claude-scholar5.7k1 repos~2.4kAutomated safety check: PassMIT
Bio Metagenomics VisualizationGPTomics/bioSkills1.2k1 repos~3.7kAutomated safety check: PassMIT
Research Analysis Routerwentorai/Research-Claw858—~169Automated safety check: PassCustom licence
Analysis Graphingclshortfuse/renodx4.5k—~1.1kAutomated safety check: PassMIT
CSV Data Analysis5zjk5/prompt-engineering127—~2.6kAutomated safety check: PassNone
Experiment Results Analysis for PapersLigphiDonk/Oh-my--paper739—~3kAutomated safety check: PassMIT

Similar skills

  • Turns a shotgun profiler table (MetaPhlAn relative abundance, Bracken counts, HUMAnN function tables) into honest figures and defensible community statistics with phyloseq, vegan, microViz, and…

    1.2k GitHub starsUsed in 1 repo~3.7k tokens
    Data & AnalyticsAuto-check passed
  • Research Analysis Router

    wentorai/Research-Claw

    科研分析能力入口:统计、因果推断、数据清洗与可视化。Use for statistics, econometrics, wrangling, causal inference, charts, and publication figures.

    858 GitHub stars~169 tokensUpdated 1 mo ago
    Data & AnalyticsAuto-check passed
  • Analysis Graphing

    clshortfuse/renodx

    RenoDX workflow for creating readable analysis graphs and plots from shader math, CSVs, EXRs, LUTs, hue sweeps, tone curves, gamut comparisons, energy/scalar maps, and test-pattern statistics.

    4.5k GitHub stars~1.1k tokensUpdated yesterday
    Data & AnalyticsAuto-check passed
  • CSV Data Analysis

    5zjk5/prompt-engineering

    This skill should be used when users need to analyze CSV or Excel files, understand data patterns, generate statistical summaries, or create data visualizations.

    127 GitHub stars~2.6k tokensUpdated 25 days ago
    Data & AnalyticsAuto-check passed
  • Turns experimental data such as CSV, JSON or TensorBoard logs into statistical significance tests, visualizations and a drafted Results section.

    739 GitHub stars~3k tokensUpdated 5 mo ago
    Data & AnalyticsAuto-check passed
  • Data Analysis

    fastclaw-ai/fastclaw

    Analyze data, process CSV/JSON files, compute statistics, and create data visualizations.

    1.4k GitHub stars~410 tokensUpdated yesterday
    Data & AnalyticsAuto-check passed

More from Galaxy-Dawn/claude-scholar

All 34 skills in this repo
  • Citation Verification Guide

    Galaxy-Dawn/claude-scholar

    Reference guidance for checking every citation in academic writing against canonical sources such as DOI, arXiv, CrossRef and Semantic Scholar, to catch fake or wrong references.

    5.7k GitHub starsUsed in 2 repos~1.9k tokens
    Auto-check passed
  • Skill Improvement Plan Executor

    Galaxy-Dawn/claude-scholar

    Reads an improvement-plan file from a companion quality-review skill and applies its suggested fixes to a Claude Skill, backing up first.

    5.7k GitHub starsUsed in 1 repo~1.8k tokens
    Auto-check passed
  • Skill Quality Reviewer

    Galaxy-Dawn/claude-scholar

    Scores a skill across description, content organization, writing style and structure, then produces letter grades and a prioritized improvement plan.

    5.7k GitHub starsUsed in 1 repo~3k tokens
    Auto-check passed
  • UI/UX Design System Advisor

    Galaxy-Dawn/claude-scholar

    Turns a vague UI request into a concrete design system with style, palette, typography and layout guidance from a search script, plus stack-specific implementation advice.

    5.7k GitHub starsUsed in 1 repo~1.1k tokens
    Auto-check passed
  • Daily Paper Digest

    Galaxy-Dawn/claude-scholar

    Finds recent arXiv and bioRxiv papers on a topic, narrows them in stages to one pick per field, and writes bilingual Chinese and English summaries.

    5.7k GitHub stars~1k tokensUpdated 17 days ago
    Auto-check passed
  • Nature Data

    Galaxy-Dawn/claude-scholar

    Prepare, audit, or revise Nature-ready Data Availability statements, data repository plans, dataset citations, and FAIR metadata checklists for manuscripts.

    5.7k GitHub starsUsed in 3 repos~1.6k tokens
    Auto-check passed

Questions about Results Analysis

What does Results Analysis do?

This skill should be used when the user asks to "analyze experimental results", "run strict statistical analysis", "compare model performance", "generate scientific figures", "check significance"…. Results Analysis is an agent skill from Galaxy-Dawn/claude-scholar. This skill should be used when the user asks to "analyze experimental results", "run strict statistical analysis", "compare model performance", "generate scientific figures", "check significance", "do ablation analysis", or mentions interpreting experiment data with rigorous statistics and visualization.

When should I use Results Analysis?

Results Analysis fits situations like: asks to analyze experimental results; run strict statistical analysis; compare model performance; generate scientific figures.

How do I install Results Analysis in Claude Code?

Run `npx skills add Galaxy-Dawn/claude-scholar --skill results-analysis -a claude-code`. Or copy the skill folder (skills/results-analysis in Galaxy-Dawn/claude-scholar) into .claude/skills/results-analysis in your project. Claude Code loads it when a task matches its description.

How do I install Results Analysis in Codex?

Run `npx skills add Galaxy-Dawn/claude-scholar --skill results-analysis -a codex`. Or copy the skill folder (skills/results-analysis in Galaxy-Dawn/claude-scholar) into .agents/skills/results-analysis in your project. Codex loads it when a task matches its description.

Can I use Results Analysis in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Galaxy-Dawn/claude-scholar --skill results-analysis -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/results-analysis, .gemini/skills/results-analysis, .github/skills/results-analysis and .opencode/skills/results-analysis in your project.

What does Results Analysis need to run?

SKILL.md names no scripts, command-line tools or credentials: Results Analysis is instructions for the agent only.

Does Results Analysis access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Results Analysis safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Results Analysis use?

Results Analysis is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Results Analysis use?

About 2.4k tokens (SKILL.md is roughly 9.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 6.7k tokens, read only when the agent opens those files.

What are the alternatives to Results Analysis?

Skills that share tags, products or a category with Results Analysis: Bio Metagenomics Visualization (GPTomics/bioSkills, 1.2k stars), Research Analysis Router (wentorai/Research-Claw, 858 stars), Analysis Graphing (clshortfuse/renodx, 4.5k stars) and CSV Data Analysis (5zjk5/prompt-engineering, 127 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Results Analysis?

Galaxy-Dawn (a GitHub user) maintains it in Galaxy-Dawn/claude-scholar, which has 5,725 GitHub stars. The repository holds 34 skills in this directory. The repository was last updated on September 23, 2026.

Source: Galaxy-Dawn/claude-scholar on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.