Agent skill

Python Reproducibility Guide

by wentorai in wentorai/research-plugins

Reproducible Python environments, notebooks, and literate programming

MITAuto-check passedResearch & Science

Install Python Reproducibility Guide

skills CLI
$ npx skills add wentorai/research-plugins --skill python-reproducibility-guide -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install wentorai/research-plugins python-reproducibility-guide --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/wentorai/research-plugins.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/tools/code-exec/python-reproducibility-guide .claude/skills/python-reproducibility-guide && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
python-reproducibility-guide
GitHub stars
298
Used in
1 other repo
Token cost
~2.2k tokens
SKILL.md length
157 words
Files
1
Skills in repo
428
Repo updated
First seen
Licence
MIT

At a glance

Reproducible Python environments, notebooks, and literate programming

  • Tasks that involve Reproducible research
  • SKILL.md covers Environment Management, Jupyter Notebooks for Research, Reproducible Random Seeds and Containerization with Docker, plus 4 more sections
  • Calls jupyter, conda and pip
  • Tasks that involve Containers

What it does

Python Reproducibility Guide is an agent skill from wentorai/research-plugins. Reproducible Python environments, notebooks, and literate programming

Its SKILL.md is about 2.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Research & Science, covering Reproducible research and Containers. It works with Python. The repository describes itself as: 350+ academic research skills, MCP configs, and plugins for Research-Claw and AI agents. The licence is MIT.

When your agent uses it

  • Tasks that involve Reproducible research
  • Tasks that involve Containers

Example prompts

  • “/python-reproducibility-guide”

Requirements

  • Python 3
  • Docker

What it can do on your machine

Read from SKILL.md and the folder at commit bf44b3c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • jupyter
    • conda
    • pip
    • docker
    • python
    • uv
    • make

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, docker and uv, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Python Reproducibility Guide loads about 2.2k tokens when it runs. Until then it costs about 25 tokens; SKILL.md has 157 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~25
When it runs · the whole SKILL.md, loaded when a task matches
~2.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from wentorai/research-plugins at commit bf44b3c, republished under its MIT licence (© wentorai). 157 words, ~2,165 tokens.

Download SKILL.mdSave it as .claude/skills/python-reproducibility-guide/SKILL.md (or your agent's skills folder).
name
python-reproducibility-guide
description
Reproducible Python environments, notebooks, and literate programming

Python Reproducibility Guide

Set up reproducible Python environments for research computing, using virtual environments, dependency management, Jupyter notebooks, and literate programming practices.

Environment Management

Virtual Environments
bash
# Option 1: venv (built-in, lightweight)
python -m venv .venv
source .venv/bin/activate       # macOS/Linux
# .venv\Scripts\activate        # Windows
pip install -r requirements.txt

# Option 2: conda (includes non-Python dependencies)
conda create -n myproject python=3.11
conda activate myproject
conda install numpy pandas scipy matplotlib
conda env export > environment.yml

# Option 3: uv (fast, modern Python package manager)
uv venv
source .venv/bin/activate
uv pip install -r requirements.txt
Dependency Pinning
bash
# requirements.txt with exact versions (pip freeze)
pip freeze > requirements.txt

# Better: use pip-tools for compiled dependencies
pip install pip-tools

# Create requirements.in (human-readable, loose constraints)
cat > requirements.in << 'EOF'
numpy>=1.24
pandas>=2.0
scipy>=1.11
matplotlib>=3.7
scikit-learn>=1.3
EOF

# Compile to requirements.txt (pinned, reproducible)
pip-compile requirements.in --output-file requirements.txt

# Install from compiled requirements
pip-sync requirements.txt
pyproject.toml (Modern Standard)
toml
[project]
name = "my-research-project"
version = "0.1.0"
description = "Analysis code for paper: Title"
requires-python = ">=3.10"
dependencies = [
    "numpy>=1.24",
    "pandas>=2.0",
    "scipy>=1.11",
    "matplotlib>=3.7",
    "scikit-learn>=1.3",
    "statsmodels>=0.14",
]

[project.optional-dependencies]
dev = ["pytest", "black", "ruff", "jupyter"]
gpu = ["torch>=2.0", "torchvision"]

[tool.ruff]
line-length = 88
select = ["E", "F", "I"]

Jupyter Notebooks for Research

Best Practices
python
# Cell 1: Imports and configuration (always the first cell)
import numpy as np
import pandas as pd
import matplotlib.pyplot as plt
from pathlib import Path

# Configuration
DATA_DIR = Path("./data")
OUTPUT_DIR = Path("./outputs")
OUTPUT_DIR.mkdir(exist_ok=True)

RANDOM_SEED = 42
np.random.seed(RANDOM_SEED)

# Matplotlib defaults
plt.rcParams.update({
    "figure.figsize": (10, 6),
    "figure.dpi": 150,
    "font.size": 12,
    "axes.spines.top": False,
    "axes.spines.right": False,
})

print(f"NumPy: {np.__version__}")
print(f"Pandas: {pd.__version__}")
Notebook Structure Template
markdown
# Paper Title: Analysis Notebook

## 1. Setup and Data Loading
[Import libraries, set seeds, load data]

## 2. Data Exploration
[Summary statistics, distributions, missing data check]

## 3. Preprocessing
[Cleaning, transformation, feature engineering]

## 4. Analysis
### 4.1 Primary Analysis
[Main statistical tests or model training]
### 4.2 Sensitivity Analysis
[Robustness checks]
### 4.3 Supplementary Analysis
[Additional analyses for appendix]

## 5. Visualization
[Publication-quality figures]

## 6. Export Results
[Save tables, figures, and summary statistics]
Converting Notebooks to Scripts
bash
# Convert notebook to Python script
jupyter nbconvert --to script analysis.ipynb

# Convert notebook to HTML report
jupyter nbconvert --to html --no-input analysis.ipynb

# Convert notebook to PDF
jupyter nbconvert --to pdf analysis.ipynb

# Execute notebook from command line (and save output)
jupyter nbconvert --execute --to notebook --inplace analysis.ipynb

Reproducible Random Seeds

python
import numpy as np
import random
import os

def set_global_seed(seed=42):
    """Set random seeds for full reproducibility."""
    random.seed(seed)
    np.random.seed(seed)
    os.environ["PYTHONHASHSEED"] = str(seed)

    # PyTorch (if used)
    try:
        import torch
        torch.manual_seed(seed)
        torch.cuda.manual_seed_all(seed)
        torch.backends.cudnn.deterministic = True
        torch.backends.cudnn.benchmark = False
    except ImportError:
        pass

    # TensorFlow (if used)
    try:
        import tensorflow as tf
        tf.random.set_seed(seed)
    except ImportError:
        pass

set_global_seed(42)

Containerization with Docker

Dockerfile for Research
dockerfile
FROM python:3.11-slim

WORKDIR /app

# System dependencies
RUN apt-get update && apt-get install -y \
    build-essential \
    git \
    && rm -rf /var/lib/apt/lists/*

# Python dependencies
COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt

# Copy project code
COPY . .

# Default: run the analysis
CMD ["python", "run_analysis.py"]
bash
# Build and run
docker build -t my-analysis .
docker run -v $(pwd)/data:/app/data -v $(pwd)/outputs:/app/outputs my-analysis

# Interactive Jupyter inside Docker
docker run -p 8888:8888 -v $(pwd):/app my-analysis \
    jupyter notebook --ip=0.0.0.0 --allow-root --no-browser

Project Structure

research-project/
├── README.md                 # Project overview and how to reproduce
├── pyproject.toml            # Dependencies and project metadata
├── requirements.txt          # Pinned dependencies
├── Dockerfile                # Containerized environment
├── Makefile                  # Automation (make data, make analysis, make figures)
├── data/
│   ├── raw/                  # Original, immutable data
│   ├── processed/            # Cleaned, transformed data
│   └── external/             # Third-party data sources
├── notebooks/
│   ├── 01_exploration.ipynb  # Data exploration
│   ├── 02_analysis.ipynb     # Main analysis
│   └── 03_figures.ipynb      # Publication figures
├── src/
│   ├── __init__.py
│   ├── data.py               # Data loading and preprocessing
│   ├── models.py             # Statistical models and ML
│   ├── visualization.py      # Plotting functions
│   └── utils.py              # Shared utilities
├── tests/
│   ├── test_data.py          # Data pipeline tests
│   └── test_models.py        # Model correctness tests
├── outputs/
│   ├── figures/              # Generated figures (PDF, PNG)
│   ├── tables/               # Generated tables (CSV, LaTeX)
│   └── models/               # Saved model artifacts
└── configs/
    ├── experiment_1.yaml     # Experiment configuration
    └── experiment_2.yaml     # Experiment configuration

Makefile for Automation

makefile
.PHONY: all data analysis figures clean

all: data analysis figures

data:
	python src/data.py --input data/raw/ --output data/processed/

analysis: data
	python -m jupyter nbconvert --execute notebooks/02_analysis.ipynb \
		--to notebook --inplace

figures: analysis
	python src/visualization.py --output outputs/figures/

clean:
	rm -rf data/processed/ outputs/

# Reproduce the full pipeline from scratch
reproduce: clean all
	@echo "All results reproduced successfully."

# Run tests
test:
	pytest tests/ -v

# Format code
format:
	ruff check --fix src/ tests/
	ruff format src/ tests/

Logging and Experiment Tracking

python
import logging
from datetime import datetime

# Set up logging
logging.basicConfig(
    level=logging.INFO,
    format="%(asctime)s [%(levelname)s] %(message)s",
    handlers=[
        logging.FileHandler(f"outputs/logs/run_{datetime.now():%Y%m%d_%H%M%S}.log"),
        logging.StreamHandler()
    ]
)
logger = logging.getLogger(__name__)

# Log experiment parameters
logger.info(f"Random seed: {RANDOM_SEED}")
logger.info(f"Data file: {DATA_DIR / 'dataset.csv'}")
logger.info(f"Model: Linear Regression with L2 regularization (alpha=0.1)")
logger.info(f"Train/test split: 80/20")

Reproducibility Checklist

  • All dependencies are pinned in requirements.txt or pyproject.toml
  • Random seeds are set at the beginning of every script/notebook
  • Raw data is stored separately and never modified
  • Data preprocessing steps are scripted (not manual)
  • Analysis can be re-run with a single command (make all or python run_analysis.py)
  • Environment is documented (Python version, OS, hardware specs)
  • Figures are generated programmatically (not edited manually)
  • Code is tested (at least smoke tests for critical functions)
  • A README explains how to set up the environment and reproduce results
  • Version control (git) tracks all code changes with meaningful commits

© wentorai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/tools/code-exec/python-reproducibility-guide of wentorai/research-plugins.

Open the folder on GitHubat commit bf44b3c

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in wentorai/research-plugins, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Python Reproducibility Guide next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Python Reproducibility Guide compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Python Reproducibility Guide this skillwentorai/research-plugins2981 repos~2.2kAutomated safety check: PassMIT
Capture Environmentpedrohcgs/claude-code-my-workflow1.6k—~2.8kAutomated safety check: NotesMIT
Bio Workflow Management Nf Core PipelinesGPTomics/bioSkills1.2k1 repos~4.1kAutomated safety check: PassMIT
Modeling Code and Result Contractsyushui2022/MathModel-Skill452—~1.4kAutomated safety check: PassMIT
HypoGeniC Hypothesis GenerationK-Dense-AI/scientific-agent-skills48k1 repos~3.6kAutomated safety check: NotesMIT
Backward Traceabilitylingzhi227/agent-research-skills383—~802Automated safety check: PassNone

Similar skills

  • Capture Environment

    pedrohcgs/claude-code-my-workflow

    Snapshot the computational environment for a replication package — detects the analysis stack (R / Stata / Python) and emits the right lockfiles (renv.lock + sessionInfo.txt, requirements.txt /…

    1.6k GitHub stars~2.8k tokensUpdated 9 days ago
    Research & ScienceAuto-check: notes
  • Runs and configures curated nf-core community Nextflow pipelines (rnaseq, sarek, atacseq, methylseq, ampliseq, taxprofiler, fetchngs) reproducibly, pinning the pipeline revision with -r and…

    1.2k GitHub starsUsed in 1 repo~4.1k tokens
    Research & ScienceAuto-check passed
  • Modeling Code and Result Contracts

    yushui2022/MathModel-Skill

    Generates result-evidence contracts, tables and runnable q1 to q3 modeling code scaffolds for a math modeling paper from a model route, a data plan and cleaned data.

    452 GitHub stars~1.4k tokensUpdated 10 days ago
    Research & ScienceAuto-check passed
  • HypoGeniC Hypothesis Generation

    K-Dense-AI/scientific-agent-skills

    Plans and audits runs of the HypoGeniC and HypoRefine packages, which propose hypotheses from labeled text datasets, with local checks before any model call.

    48k GitHub starsUsed in 1 repo~3.6k tokens
    Research & ScienceAuto-check: notes
  • Backward Traceability

    lingzhi227/agent-research-skills

    Makes each number in a LaTeX paper link back to the code line that produced it, using hypertarget and hyperlink tags and compile-time `\num` formulas.

    383 GitHub stars~802 tokensUpdated 7 mo ago
    Research & ScienceAuto-check passed
  • Literature Review

    K-Dense-AI/scientific-agent-skills

    Runs systematic, scoping or narrative literature reviews across PubMed, arXiv, bioRxiv and Semantic Scholar, with citation checks and Markdown or PDF output.

    48k GitHub starsUsed in 1 repo~3.2k tokens
    Research & ScienceAuto-check: notes

More from wentorai/research-plugins

All 428 skills in this repo
  • Abstract Writing Guide

    wentorai/research-plugins

    Craft structured research abstracts that maximize clarity and journal acceptance

    298 GitHub starsUsed in 1 repo~1.7k tokens
    Auto-check passed
  • Academic Citation Manager

    wentorai/research-plugins

    Manage academic citations across BibTeX, APA, MLA, and Chicago formats

    298 GitHub starsUsed in 1 repo~2.7k tokens
    Auto-check passed
  • Academic Paper Summarizer

    wentorai/research-plugins

    Summarize academic papers with structured extraction of key elements

    298 GitHub starsUsed in 1 repo~1.4k tokens
    Auto-check passed
  • Academic Study Methods

    wentorai/research-plugins

    Evidence-based study techniques for academic learning and retention

    298 GitHub starsUsed in 1 repo~1.8k tokens
    Auto-check passed
  • Academic Tone Guide

    wentorai/research-plugins

    Adjust writing tone and register for academic audiences and venues

    298 GitHub starsUsed in 1 repo~1.9k tokens
    Auto-check passed
  • Academic Translation Guide

    wentorai/research-plugins

    Academic translation, post-editing, and Chinglish correction guide

    298 GitHub starsUsed in 1 repo~1.6k tokens
    Auto-check passed

Works with

Questions about Python Reproducibility Guide

What does Python Reproducibility Guide do?

Reproducible Python environments, notebooks, and literate programming. Python Reproducibility Guide is an agent skill from wentorai/research-plugins.

When should I use Python Reproducibility Guide?

Python Reproducibility Guide fits situations like: tasks that involve Reproducible research; tasks that involve Containers.

How do I install Python Reproducibility Guide in Claude Code?

Run `npx skills add wentorai/research-plugins --skill python-reproducibility-guide -a claude-code`. Or copy the skill folder (skills/tools/code-exec/python-reproducibility-guide in wentorai/research-plugins) into .claude/skills/python-reproducibility-guide in your project. Claude Code loads it when a task matches its description.

How do I install Python Reproducibility Guide in Codex?

Run `npx skills add wentorai/research-plugins --skill python-reproducibility-guide -a codex`. Or copy the skill folder (skills/tools/code-exec/python-reproducibility-guide in wentorai/research-plugins) into .agents/skills/python-reproducibility-guide in your project. Codex loads it when a task matches its description.

Can I use Python Reproducibility Guide in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add wentorai/research-plugins --skill python-reproducibility-guide -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/python-reproducibility-guide, .gemini/skills/python-reproducibility-guide, .github/skills/python-reproducibility-guide and .opencode/skills/python-reproducibility-guide in your project.

What does Python Reproducibility Guide need to run?

Going by SKILL.md and its folder, Python Reproducibility Guide needs the command-line tools its instructions call (jupyter, conda, pip, docker, python and uv). Our summary lists: Python 3; Docker.

Does Python Reproducibility Guide access the network?

SKILL.md contains no URLs. Its commands use pip, docker and uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Python Reproducibility Guide safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Python Reproducibility Guide use?

Python Reproducibility Guide is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Python Reproducibility Guide use?

About 2.2k tokens (SKILL.md is roughly 8.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Python Reproducibility Guide?

Skills that share tags, products or a category with Python Reproducibility Guide: Capture Environment (pedrohcgs/claude-code-my-workflow, 1.6k stars), Bio Workflow Management Nf Core Pipelines (GPTomics/bioSkills, 1.2k stars), Modeling Code and Result Contracts (yushui2022/MathModel-Skill, 452 stars) and HypoGeniC Hypothesis Generation (K-Dense-AI/scientific-agent-skills, 48k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Python Reproducibility Guide?

wentorai (a GitHub user) maintains it in wentorai/research-plugins, which has 298 GitHub stars. The repository holds 428 skills in this directory. The repository was last updated on June 19, 2026.

Source: wentorai/research-plugins on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.