Agent skill

Data Stats Analysis

by aipoch in aipoch/medical-research-skills

Perform statistical tests, hypothesis testing, correlation analysis, and multiple testing corrections using scipy and statsmodels.

MITAuto-check passedData & Analytics

Install Data Stats Analysis

skills CLI
$ npx skills add aipoch/medical-research-skills --skill data-stats-analysis -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install aipoch/medical-research-skills data-stats-analysis --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/aipoch/medical-research-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/'scientific-skills/Data Analysis/data-stats-analysis' .claude/skills/data-stats-analysis && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
data-stats-analysis
GitHub stars
2k
Used in
2 other repos
Token cost
~4k tokens
SKILL.md length
474 words
Files
3
Skills in repo
578
Repo updated
First seen
Licence
MIT

At a glance

Perform statistical tests, hypothesis testing, correlation analysis, and multiple testing corrections using scipy and statsmodels.

  • Works in 6 steps: Import Required Libraries → Two-Sample t-Test → One-Way ANOVA → …
  • Tasks that involve Statistics
  • SKILL.md covers Overview, When to Use This Skill, How to Use and Advanced Features, plus 6 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Data Stats Analysis is an agent skill from aipoch/medical-research-skills. Perform statistical tests, hypothesis testing, correlation analysis, and multiple testing corrections using scipy and statsmodels. Works with ANY LLM provider (GPT, Gemini, Claude, etc.).

Its SKILL.md is about 4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `POLISH_CHANGELOG.md` and `eval_report_data-stats-analysis_result.json`).

It sits in Data & Analytics, covering Statistics. It works with statsmodels. The repository describes itself as: Hundreds of agent skills for medical research, including protocol design, data analysis, evidence insights, and academic writing. The licence is MIT.

When your agent uses it

  • Tasks that involve Statistics

Example prompts

  • “/data-stats-analysis”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Import Required Libraries
  2. Two-Sample t-Test
  3. One-Way ANOVA
  4. Correlation Analysis
  5. Multiple Testing Correction
  6. Non-Parametric Tests

What it can do on your machine

Read from SKILL.md and the folder at commit 686e09d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • docs.scipy.org
    • statsmodels.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Data Stats Analysis loads about 4k tokens when it runs. Until then it costs about 52 tokens; SKILL.md has 474 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~52
When it runs · the whole SKILL.md, loaded when a task matches
~4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from aipoch/medical-research-skills at commit 686e09d, republished under its MIT licence (© aipoch). 474 words, ~3,993 tokens.

Download SKILL.mdSave it as .claude/skills/data-stats-analysis/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
data-stats-analysis
description
Perform statistical tests, hypothesis testing, correlation analysis, and multiple testing corrections using scipy and statsmodels. Works with ANY LLM provider (GPT, Gemini, Claude, etc.).
license
MIT
author
AIPOCH

Source: https://github.com/aipoch/medical-research-skills

Statistical Analysis (Universal)

Overview

This skill enables you to perform rigorous statistical analyses including t-tests, ANOVA, correlation analysis, hypothesis testing, and multiple testing corrections. Unlike cloud-hosted solutions, this skill uses standard Python statistical libraries (scipy, statsmodels, numpy) and executes locally in your environment, making it compatible with ALL LLM providers including GPT, Gemini, Claude, DeepSeek, and Qwen.

When to Use This Skill

  • Compare means between groups (t-tests, ANOVA)
  • Test for correlations between variables
  • Perform hypothesis testing with p-value calculation
  • Apply multiple testing corrections (FDR, Bonferroni)
  • Calculate statistical summaries and confidence intervals
  • Test for normality and distribution fitting
  • Perform non-parametric tests (Mann-Whitney, Kruskal-Wallis)

How to Use

Step 1: Import Required Libraries
python
import numpy as np
import pandas as pd
from scipy import stats
from scipy.stats import ttest_ind, mannwhitneyu, pearsonr, spearmanr
from scipy.stats import f_oneway, kruskal, chi2_contingency
from statsmodels.stats.multitest import multipletests
from statsmodels.stats.proportion import proportions_ztest
import warnings
warnings.filterwarnings('ignore')
Step 2: Two-Sample t-Test
python
# Compare means between two groups
# group1, group2: arrays of numeric values

# Perform independent t-test
t_statistic, p_value = ttest_ind(group1, group2)

print(f"t-statistic: {t_statistic:.4f}")
print(f"p-value: {p_value:.4e}")

if p_value < 0.05:
    print("✅ Significant difference between groups (p < 0.05)")
else:
    print("❌ No significant difference (p >= 0.05)")

# With equal variance assumption check
# Levene's test for equal variances
_, levene_p = stats.levene(group1, group2)
if levene_p < 0.05:
    # Use Welch's t-test (unequal variances)
    t_stat, p_val = ttest_ind(group1, group2, equal_var=False)
    print(f"Welch's t-test p-value: {p_val:.4e}")
else:
    print("Equal variances assumed")
Step 3: One-Way ANOVA
python
# Compare means across multiple groups
# groups: list of arrays, e.g., [group1, group2, group3]

# Perform one-way ANOVA
f_statistic, p_value = f_oneway(*groups)

print(f"F-statistic: {f_statistic:.4f}")
print(f"p-value: {p_value:.4e}")

if p_value < 0.05:
    print("✅ Significant difference between groups (p < 0.05)")
    print("Note: Use post-hoc tests to identify which groups differ")
else:
    print("❌ No significant difference between groups")

# Post-hoc pairwise t-tests with Bonferroni correction
from itertools import combinations

group_names = ['Group A', 'Group B', 'Group C']
pairwise_results = []

for (name1, data1), (name2, data2) in combinations(zip(group_names, groups), 2):
    _, p = ttest_ind(data1, data2)
    pairwise_results.append({
        'comparison': f'{name1} vs {name2}',
        'p_value': p
    })

# Apply Bonferroni correction
pairwise_df = pd.DataFrame(pairwise_results)
n_tests = len(pairwise_df)
pairwise_df['p_adjusted'] = pairwise_df['p_value'] * n_tests
pairwise_df['p_adjusted'] = pairwise_df['p_adjusted'].clip(upper=1.0)

print("\nPairwise Comparisons (Bonferroni-corrected):")
print(pairwise_df)
Step 4: Correlation Analysis
python
# Pearson correlation (linear relationships)
r_pearson, p_pearson = pearsonr(variable1, variable2)

print(f"Pearson correlation: r = {r_pearson:.4f}, p = {p_pearson:.4e}")

# Spearman correlation (monotonic relationships, robust to outliers)
r_spearman, p_spearman = spearmanr(variable1, variable2)

print(f"Spearman correlation: ρ = {r_spearman:.4f}, p = {p_spearman:.4e}")

# Interpretation
if abs(r_pearson) < 0.3:
    strength = "weak"
elif abs(r_pearson) < 0.7:
    strength = "moderate"
else:
    strength = "strong"

direction = "positive" if r_pearson > 0 else "negative"
print(f"Interpretation: {strength} {direction} correlation")

if p_pearson < 0.05:
    print("✅ Statistically significant (p < 0.05)")
else:
    print("❌ Not statistically significant")
Step 5: Multiple Testing Correction
python
# Scenario: Testing 1000 genes for differential expression
# p_values: array of p-values from individual tests

# Method 1: Benjamini-Hochberg FDR correction (recommended)
reject_fdr, p_adjusted_fdr, _, _ = multipletests(p_values, alpha=0.05, method='fdr_bh')

# Method 2: Bonferroni correction (more conservative)
reject_bonf, p_adjusted_bonf, _, _ = multipletests(p_values, alpha=0.05, method='bonferroni')

# Create results DataFrame
results_df = pd.DataFrame({
    'gene': gene_names,
    'p_value': p_values,
    'q_value_fdr': p_adjusted_fdr,
    'p_adjusted_bonferroni': p_adjusted_bonf,
    'significant_fdr': reject_fdr,
    'significant_bonf': reject_bonf
})

# Summary
print(f"Original significant (p < 0.05): {(p_values < 0.05).sum()}")
print(f"Significant after FDR correction: {reject_fdr.sum()}")
print(f"Significant after Bonferroni correction: {reject_bonf.sum()}")

# Save results
results_df.to_csv('statistical_results.csv', index=False)
print("✅ Results saved to: statistical_results.csv")
Step 6: Non-Parametric Tests
python
# Use when data is not normally distributed

# Mann-Whitney U test (alternative to t-test)
u_statistic, p_value_mw = mannwhitneyu(group1, group2, alternative='two-sided')

print(f"Mann-Whitney U test:")
print(f"U-statistic: {u_statistic:.4f}")
print(f"p-value: {p_value_mw:.4e}")

# Kruskal-Wallis H test (alternative to ANOVA)
h_statistic, p_value_kw = kruskal(*groups)

print(f"\nKruskal-Wallis H test:")
print(f"H-statistic: {h_statistic:.4f}")
print(f"p-value: {p_value_kw:.4e}")

Advanced Features

Normality Testing
python
from scipy.stats import shapiro, normaltest, kstest

# Test if data follows normal distribution

# Shapiro-Wilk test (best for n < 5000)
stat_sw, p_sw = shapiro(data)
print(f"Shapiro-Wilk test: W={stat_sw:.4f}, p={p_sw:.4e}")

# D'Agostino-Pearson test
stat_dp, p_dp = normaltest(data)
print(f"D'Agostino-Pearson test: stat={stat_dp:.4f}, p={p_dp:.4e}")

# Interpretation
if p_sw < 0.05:
    print("❌ Data does NOT follow normal distribution (p < 0.05)")
    print("→ Recommendation: Use non-parametric tests (Mann-Whitney, Kruskal-Wallis)")
else:
    print("✅ Data appears normally distributed (p >= 0.05)")
    print("→ OK to use parametric tests (t-test, ANOVA)")
Chi-Square Test for Contingency Tables
python
# Test independence between categorical variables
# contingency_table: 2D array (rows=categories1, columns=categories2)

# Example: Cell type distribution across conditions
contingency_table = np.array([
    [50, 30, 20],  # Condition A: T cells, B cells, NK cells
    [40, 45, 15],  # Condition B
    [35, 25, 40]   # Condition C
])

chi2, p_value, dof, expected = chi2_contingency(contingency_table)

print(f"Chi-square statistic: {chi2:.4f}")
print(f"p-value: {p_value:.4e}")
print(f"Degrees of freedom: {dof}")
print(f"\nExpected frequencies:\n{expected}")

if p_value < 0.05:
    print("✅ Significant association between variables (p < 0.05)")
else:
    print("❌ No significant association")
Confidence Intervals
python
from scipy.stats import t as t_dist

def calculate_confidence_interval(data, confidence=0.95):
    """Calculate confidence interval for mean"""
    n = len(data)
    mean = np.mean(data)
    std_err = stats.sem(data)  # Standard error of mean

    # t-distribution critical value
    t_crit = t_dist.ppf((1 + confidence) / 2, df=n-1)

    margin_error = t_crit * std_err
    ci_lower = mean - margin_error
    ci_upper = mean + margin_error

    return mean, ci_lower, ci_upper

# Usage
mean, ci_low, ci_high = calculate_confidence_interval(data, confidence=0.95)

print(f"Mean: {mean:.4f}")
print(f"95% CI: [{ci_low:.4f}, {ci_high:.4f}]")
Effect Size Calculation
python
def cohens_d(group1, group2):
    """Calculate Cohen's d effect size"""
    n1, n2 = len(group1), len(group2)
    var1, var2 = np.var(group1, ddof=1), np.var(group2, ddof=1)

    # Pooled standard deviation
    pooled_std = np.sqrt(((n1-1)*var1 + (n2-1)*var2) / (n1+n2-2))

    # Cohen's d
    d = (np.mean(group1) - np.mean(group2)) / pooled_std

    return d

# Usage
effect_size = cohens_d(group1, group2)
print(f"Cohen's d: {effect_size:.4f}")

# Interpretation
if abs(effect_size) < 0.2:
    print("Effect size: negligible")
elif abs(effect_size) < 0.5:
    print("Effect size: small")
elif abs(effect_size) < 0.8:
    print("Effect size: medium")
else:
    print("Effect size: large")

Common Use Cases

Differential Gene Expression Statistical Testing
python
# Compare gene expression between two conditions
# gene_expression_df: rows=genes, columns=samples
# condition_labels: array indicating which condition each sample belongs to

results = []

for gene in gene_expression_df.index:
    # Get expression values for each condition
    cond1_expr = gene_expression_df.loc[gene, condition_labels == 'Condition1']
    cond2_expr = gene_expression_df.loc[gene, condition_labels == 'Condition2']

    # t-test
    t_stat, p_val = ttest_ind(cond1_expr, cond2_expr)

    # Log2 fold change
    log2fc = np.log2(cond2_expr.mean() / cond1_expr.mean())

    results.append({
        'gene': gene,
        'log2FC': log2fc,
        'p_value': p_val,
        'mean_cond1': cond1_expr.mean(),
        'mean_cond2': cond2_expr.mean()
    })

deg_results = pd.DataFrame(results)

# Apply FDR correction
_, deg_results['q_value'], _, _ = multipletests(
    deg_results['p_value'],
    alpha=0.05,
    method='fdr_bh'
)

# Filter significant genes
significant_genes = deg_results[
    (deg_results['q_value'] < 0.05) &
    (abs(deg_results['log2FC']) > 1)
]

print(f"✅ Identified {len(significant_genes)} differentially expressed genes")
print(f"   - Upregulated: {(significant_genes['log2FC'] > 1).sum()}")
print(f"   - Downregulated: {(significant_genes['log2FC'] < -1).sum()}")

# Save
significant_genes.to_csv('deg_results.csv', index=False)
Cluster Enrichment Analysis
python
# Test if a cell type is enriched in a specific cluster
# total_cells: total number of cells
# cluster_cells: number of cells in cluster
# celltype_total: total cells of this type
# celltype_in_cluster: cells of this type in cluster

from scipy.stats import fisher_exact

# Create contingency table
contingency = [
    [celltype_in_cluster, cluster_cells - celltype_in_cluster],  # In cluster
    [celltype_total - celltype_in_cluster, total_cells - cluster_cells - (celltype_total - celltype_in_cluster)]  # Not in cluster
]

odds_ratio, p_value = fisher_exact(contingency, alternative='greater')

print(f"Odds ratio: {odds_ratio:.4f}")
print(f"p-value: {p_value:.4e}")

if p_value < 0.05 and odds_ratio > 1:
    print(f"✅ Cell type is significantly ENRICHED in cluster (p < 0.05)")
elif p_value < 0.05 and odds_ratio < 1:
    print(f"⚠️ Cell type is significantly DEPLETED in cluster (p < 0.05)")
else:
    print("❌ No significant enrichment/depletion")
Batch Effect Detection
python
# Test if there's a batch effect using ANOVA
# gene_expression: DataFrame with genes as rows, samples as columns
# batch_labels: array indicating batch for each sample

batch_effect_results = []

for gene in gene_expression.index:
    # Get expression values for each batch
    batches = [
        gene_expression.loc[gene, batch_labels == batch]
        for batch in np.unique(batch_labels)
    ]

    # ANOVA test
    f_stat, p_val = f_oneway(*batches)

    batch_effect_results.append({
        'gene': gene,
        'f_statistic': f_stat,
        'p_value': p_val
    })

batch_df = pd.DataFrame(batch_effect_results)

# Apply FDR correction
_, batch_df['q_value'], _, _ = multipletests(batch_df['p_value'], alpha=0.05, method='fdr_bh')

# Count genes with batch effects
genes_with_batch_effect = (batch_df['q_value'] < 0.05).sum()

print(f"Genes with significant batch effect: {genes_with_batch_effect} ({genes_with_batch_effect/len(batch_df)*100:.1f}%)")

if genes_with_batch_effect > len(batch_df) * 0.1:
    print("⚠️ WARNING: Strong batch effect detected (>10% genes affected)")
    print("→ Recommendation: Apply batch correction (ComBat, Harmony, etc.)")
else:
    print("✅ Minimal batch effect")

Best Practices

  1. Check Assumptions: Always test normality before using parametric tests (t-test, ANOVA)
  2. Multiple Testing: Apply FDR or Bonferroni correction when testing many hypotheses
  3. Effect Size: Report effect sizes (Cohen's d) alongside p-values
  4. Sample Size: Ensure adequate sample size for statistical power
  5. Outliers: Check for and handle outliers appropriately
  6. Non-Parametric Alternatives: Use when assumptions are violated (Mann-Whitney instead of t-test)
  7. Report Details: Always report test used, test statistic, p-value, and correction method
  8. Visualization: Combine statistical tests with visualizations (box plots, violin plots)

Troubleshooting

Issue: "Warning: p-value is very small"

Solution: This is normal for highly significant results. Report as p < 0.001 or use scientific notation

python
if p_value < 0.001:
    print(f"p < 0.001")
else:
    print(f"p = {p_value:.4f}")
Issue: "Division by zero in effect size calculation"

Solution: Check for zero variance (all values identical)

python
if np.std(group1) == 0 or np.std(group2) == 0:
    print("Cannot calculate effect size: zero variance in one or both groups")
else:
    d = cohens_d(group1, group2)
Show full SKILL.md (184 more words)Show less
Issue: "Test fails with NaN values"

Solution: Remove or impute NaN values before testing

python
# Remove NaN
group1_clean = group1[~np.isnan(group1)]
group2_clean = group2[~np.isnan(group2)]

# Or filter in DataFrame
df_clean = df.dropna(subset=['column_name'])
Issue: "Insufficient sample size warning"

Solution: Minimum sample sizes for reliable tests:

  • t-test: n ≥ 30 per group (or ≥ 5 if normally distributed)
  • ANOVA: n ≥ 20 per group
  • Correlation: n ≥ 30 total
python
if len(group1) < 30 or len(group2) < 30:
    print("⚠️ Warning: Small sample size. Results may not be reliable.")
    print("Consider using non-parametric tests or collecting more data.")

Technical Notes

  • Libraries: Uses scipy.stats and statsmodels (widely supported, stable)
  • Execution: Runs locally in the agent's sandbox
  • Compatibility: Works with ALL LLM providers (GPT, Gemini, Claude, DeepSeek, Qwen, etc.)
  • Performance: Most tests complete in milliseconds; large-scale testing (>10K genes) takes 1-5 seconds
  • Precision: Uses double-precision floating point (numpy default)
  • Corrections: FDR (Benjamini-Hochberg) recommended for genomics; Bonferroni for small numbers of tests

Input Validation

This skill accepts requests that match the documented purpose of data-stats-analysis and include enough context to complete the workflow safely.

Do not continue the workflow when the request is out of scope, missing a critical input, or would require unsupported assumptions. Instead respond:

data-stats-analysis only handles its documented workflow. Please provide the missing required inputs or switch to a more suitable skill.

References

© aipoch, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in scientific-skills/Data Analysis/data-stats-analysis of aipoch/medical-research-skills.

  • SKILL.md
  • POLISH_CHANGELOG.md
  • eval_report_data-stats-analysis_result.json

Open the folder on GitHubat commit 686e09d

Used in 2 other repositories

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in aipoch/medical-research-skills, which our catalogue first saw on October 9, 2026.

Compare with similar skills

Data Stats Analysis next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Data Stats Analysis compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Data Stats Analysis this skillaipoch/medical-research-skills2k2 repos~4kAutomated safety check: PassMIT
Statistical Analysisspacering-net/codeg3.9k3 repos~5kAutomated safety check: PassMIT
StatsmodelszLanqing/codex-claude-academic-skills4.7k15 repos~4.9kAutomated safety check: PassBSD-3-Clause
Statistical Data Analysislingzhi227/agent-research-skills386—~886Automated safety check: PassNone
Quant Statistical MethodsHKUDS/Vibe-Trading35k—~4kAutomated safety check: PassMIT
StatsmodelsK-Dense-AI/scientific-agent-skills48k1 repos~3.2kAutomated safety check: NotesBSD-3-Clause

Similar skills

  • Statistical Analysis

    spacering-net/codeg

    Guided statistical analysis for research data - test selection, assumption checking, effect sizes, power analysis, Bayesian alternatives, and APA-formatted reporting.

    3.9k GitHub starsUsed in 3 repos~5k tokens
    Data & AnalyticsAuto-check passed
  • Statsmodels

    zLanqing/codex-claude-academic-skills

    Statistical models library for Python. An agent skill from zLanqing/codex-claude-academic-skills.

    4.7k GitHub starsUsed in 15 repos~4.9k tokens
    Data & AnalyticsAuto-check passed
  • Statistical Data Analysis

    lingzhi227/agent-research-skills

    Writes statistical analysis code for experimental data, runs it through a four-round review, and reports effect sizes, p-values and confidence intervals.

    386 GitHub stars~886 tokensUpdated 7 mo ago
    Data & AnalyticsAuto-check passed
  • Quant Statistical Methods

    HKUDS/Vibe-Trading

    Guides your agent through unit-root, cointegration, GARCH, bootstrap and regression-diagnostic tests on financial time series, using a tested helper module.

    35k GitHub stars~4k tokensUpdated yesterday
    Data & AnalyticsAuto-check passed
  • Statsmodels

    K-Dense-AI/scientific-agent-skills

    Fits and diagnoses Python statistical models including OLS, GLM, discrete and mixed models, ARIMA and SARIMAX.

    48k GitHub starsUsed in 1 repo~3.2k tokens
    Data & AnalyticsAuto-check: notes
  • Linearmodels

    brycewang-stanford/Auto-Empirical-Research-Skills

    Panel data, IV/GMM, system regression. An agent skill from brycewang-stanford/Auto-Empirical-Research-Skills.

    4.5k GitHub stars~3.2k tokensUpdated 4 days ago
    Data & AnalyticsAuto-check passed

More from aipoch/medical-research-skills

All 578 skills in this repo
  • Academic Poster Generator

    aipoch/medical-research-skills

    Complete workflow for generating academic research posters from PDF literature; use when you need to extract paper content from PDFs and produce a LaTeX-based poster…

    2k GitHub stars~2.2k tokensUpdated 22 days ago
    Auto-check passed
  • Diagnostic Study Quality Assessment Quadas

    aipoch/medical-research-skills

    Analyzes clinical diagnostic accuracy studies for bias using the QUADAS-2 tool.

    2k GitHub stars~1.4k tokensUpdated 22 days ago
    Auto-check passed
  • Exploratory Data Analysis

    aipoch/medical-research-skills

    Perform comprehensive exploratory data analysis on scientific data files across 200+ file formats.

    2k GitHub stars~3.7k tokensUpdated 22 days ago
    Auto-check passed
  • Iso Certification

    aipoch/medical-research-skills

    A toolkit for preparing ISO 13485:2016 certification documentation for medical device QMS.

    2k GitHub stars~1.8k tokensUpdated 22 days ago
    Auto-check passed
  • Journal Skills

    aipoch/medical-research-skills

    Recommends target journals for manuscript submission by analyzing the paper topic/abstract and the journal distribution of similar PubMed literature; use when users ask for journal…

    2k GitHub stars~1.7k tokensUpdated 22 days ago
    Auto-check passed
  • Latex Posters

    aipoch/medical-research-skills

    Creates academic-poster writing packages for LaTeX using beamerposter, tikzposter, or baposter.

    2k GitHub stars~1.3k tokensUpdated 22 days ago
    Auto-check passed

Works with

Questions about Data Stats Analysis

What does Data Stats Analysis do?

Perform statistical tests, hypothesis testing, correlation analysis, and multiple testing corrections using scipy and statsmodels. Data Stats Analysis is an agent skill from aipoch/medical-research-skills. Perform statistical tests, hypothesis testing, correlation analysis, and multiple testing corrections using scipy and statsmodels.

When should I use Data Stats Analysis?

Data Stats Analysis fits situations like: tasks that involve Statistics.

How do I install Data Stats Analysis in Claude Code?

Run `npx skills add aipoch/medical-research-skills --skill data-stats-analysis -a claude-code`. Or copy the skill folder (scientific-skills/Data Analysis/data-stats-analysis in aipoch/medical-research-skills) into .claude/skills/data-stats-analysis in your project. Claude Code loads it when a task matches its description.

How do I install Data Stats Analysis in Codex?

Run `npx skills add aipoch/medical-research-skills --skill data-stats-analysis -a codex`. Or copy the skill folder (scientific-skills/Data Analysis/data-stats-analysis in aipoch/medical-research-skills) into .agents/skills/data-stats-analysis in your project. Codex loads it when a task matches its description.

Can I use Data Stats Analysis in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add aipoch/medical-research-skills --skill data-stats-analysis -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/data-stats-analysis, .gemini/skills/data-stats-analysis, .github/skills/data-stats-analysis and .opencode/skills/data-stats-analysis in your project.

What does Data Stats Analysis need to run?

SKILL.md names no scripts, command-line tools or credentials: Data Stats Analysis is instructions for the agent only. Our summary lists: Python 3.

Does Data Stats Analysis access the network?

SKILL.md names 2 domains. As links in the text: docs.scipy.org and statsmodels.org. This is read from the text; nothing was executed.

Is Data Stats Analysis safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Data Stats Analysis use?

Data Stats Analysis is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Data Stats Analysis use?

About 4k tokens (SKILL.md is roughly 16k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Data Stats Analysis?

Skills that share tags, products or a category with Data Stats Analysis: Statistical Analysis (spacering-net/codeg, 3.9k stars), Statsmodels (zLanqing/codex-claude-academic-skills, 4.7k stars), Statistical Data Analysis (lingzhi227/agent-research-skills, 386 stars) and Quant Statistical Methods (HKUDS/Vibe-Trading, 35k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Data Stats Analysis?

aipoch (a GitHub organization) maintains it in aipoch/medical-research-skills, which has 1,978 GitHub stars. The repository holds 578 skills in this directory. The repository was last updated on September 17, 2026.

Source: aipoch/medical-research-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.