Agent skill

Code Refactor For Reproducibility

by aipoch in aipoch/medical-research-skills

A skill your agent uses when refactoring research code for publication, adding documentation to existing analysis scripts, creating reproducible computational workflows, or preparing code for…

MITAuto-check passedDevelopment

Install Code Refactor For Reproducibility

skills CLI
$ npx skills add aipoch/medical-research-skills --skill code-refactor-for-reproducibility -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install aipoch/medical-research-skills code-refactor-for-reproducibility --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/aipoch/medical-research-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/'scientific-skills/Data Analysis/code-refactor-for-reproducibility' .claude/skills/code-refactor-for-reproducibility && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
code-refactor-for-reproducibility
GitHub stars
2k
Token cost
~3.1k tokens
SKILL.md length
946 words
Files
5 (incl. scripts)
Skills in repo
567
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when refactoring research code for publication, adding documentation to existing analysis scripts, creating reproducible computational workflows, or preparing code for…

  • Works in 4 steps: Analyze Code for Reproducibility Issues → Refactor for Best Practices → Generate Environment Specifications → …
  • Refactoring research code for publication
  • SKILL.md covers When to Use, Key Features, Dependencies and Example Usage, plus 18 more sections
  • Runs Python scripts from its folder; calls python and pytest

What it does

Code Refactor For Reproducibility is an agent skill from aipoch/medical-research-skills. Use when refactoring research code for publication, adding documentation to existing analysis scripts, creating reproducible computational workflows, or preparing code for sharing with collaborators. Transforms research code into publication-ready, reproducible workflows. Adds documentation, implements error handling, creates environment specifications, and ensures computational reproducibility for scientific publications.

Its SKILL.md is about 3.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including scripts (for example `code-refactor-for-reproducibility_audit_result_v2.json`, `scripts/main.py` and `tile.json`).

It sits in Development, covering Reproducible research and Refactoring. The repository describes itself as: Hundreds of agent skills for medical research, including protocol design, data analysis, evidence insights, and academic writing. The licence is MIT.

When your agent uses it

  • Refactoring research code for publication
  • Adding documentation to existing analysis scripts
  • Creating reproducible computational workflows
  • Preparing code for sharing with collaborators

Example prompts

  • “/code-refactor-for-reproducibility”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Analyze Code for Reproducibility Issues
  2. Refactor for Best Practices
  3. Generate Environment Specifications
  4. Create Documentation

What it can do on your machine

Read from SKILL.md and the folder at commit 686e09d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python
    • pytest

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Code Refactor For Reproducibility loads about 3.1k tokens when it runs. Until then it costs about 115 tokens; SKILL.md has 946 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~115
When it runs · the whole SKILL.md, loaded when a task matches
~3.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from aipoch/medical-research-skills at commit 686e09d, republished under its MIT licence (© aipoch). 946 words, ~3,075 tokens.

Download SKILL.mdSave it as .claude/skills/code-refactor-for-reproducibility/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
code-refactor-for-reproducibility
description
Use when refactoring research code for publication, adding documentation to existing analysis scripts, creating reproducible computational workflows, or preparing code for sharing with collaborators. Transforms research code into publication-ready, reproducible workflows. Adds documentation, implements error handling, creates environment specifications, and ensures computational reproducibility for scientific publications.
license
MIT
author
AIPOCH

Source: https://github.com/aipoch/medical-research-skills

Research Code Reproducibility Refactoring Tool

When to Use

  • Use this skill when the task needs Use when refactoring research code for publication, adding documentation to existing analysis scripts, creating reproducible computational workflows, or preparing code for sharing with collaborators. Transforms research code into publication-ready, reproducible workflows. Adds documentation, implements error handling, creates environment specifications, and ensures computational reproducibility for scientific publications.
  • Use this skill for data analysis tasks that require explicit assumptions, bounded scope, and a reproducible output format.
  • Use this skill when you need a documented fallback path for missing inputs, execution errors, or partial evidence.

Key Features

  • Scope-focused workflow aligned to: Use when refactoring research code for publication, adding documentation to existing analysis scripts, creating reproducible computational workflows, or preparing code for sharing with collaborators. Transforms research code into publication-ready, reproducible workflows. Adds documentation, implements error handling, creates environment specifications, and ensures computational reproducibility for scientific publications.
  • Packaged executable path(s): scripts/main.py.
  • Structured execution path designed to keep outputs consistent and reviewable.

Dependencies

  • Python: 3.10+. Repository baseline for current packaged skills.
  • numpy: unspecified. Declared in requirements.txt.
  • pandas: unspecified. Declared in requirements.txt.
  • pytest: unspecified. Declared in requirements.txt.
  • scipy: unspecified. Declared in requirements.txt.
  • src: unspecified. Declared in requirements.txt.

Example Usage

bash
cd "20260318/scientific-skills/Data Analytics/code-refactor-for-reproducibility"
python -m py_compile scripts/main.py
python scripts/main.py --help

Example run plan:

  1. Confirm the user input, output path, and any required config values.
  2. Edit the in-file CONFIG block or documented parameters if the script uses fixed settings.
  3. Run python scripts/main.py with the validated inputs.
  4. Review the generated output and return the final artifact with any assumptions called out.

Implementation Details

See ## Workflow above for related details.

  • Execution model: validate the request, choose the packaged workflow, and produce a bounded deliverable.
  • Input controls: confirm the source files, scope limits, output format, and acceptance criteria before running any script.
  • Primary implementation surface: scripts/main.py.
  • Parameters to clarify first: input path, output path, scope filters, thresholds, and any domain-specific constraints.
  • Output discipline: keep results reproducible, identify assumptions explicitly, and avoid undocumented side effects.

Quick Check

Use this command to verify that the packaged script entry point can be parsed before deeper execution.

bash
python -m py_compile scripts/main.py

Audit-Ready Commands

Use these concrete commands for validation. They are intentionally self-contained and avoid placeholder paths.

bash
python -m py_compile scripts/main.py
python scripts/main.py --help

Workflow

  1. Confirm the user objective, required inputs, and non-negotiable constraints before doing detailed work.
  2. Validate that the request matches the documented scope and stop early if the task would require unsupported assumptions.
  3. Use the packaged script path or the documented reasoning path with only the inputs that are actually available.
  4. Return a structured result that separates assumptions, deliverables, risks, and unresolved items.
  5. If execution fails or inputs are incomplete, switch to the fallback path and state exactly what blocked full completion.

Workflow Overview

Follow this sequence when refactoring a research codebase:

  1. Analyze — identify reproducibility issues in existing code
  2. Refactor — apply documentation, parameterization, and error handling
  3. Specify environment — pin dependencies and create environment files
  4. Validate — run tests and verify behaviour is unchanged

Step 1: Analyze Code for Reproducibility Issues

Read each source file and check for the following problems. Document findings before making any changes.

Checklist: missing docstrings · hardcoded absolute paths · missing random seeds · bare except: clauses · unpinned imports · unexplained magic numbers

Example — detecting issues manually:

python
import ast, pathlib

def find_hardcoded_paths(source: str) -> list[str]:
    """Return string literals that look like absolute paths."""
    tree = ast.parse(source)
    return [
        node.s for node in ast.walk(tree)
        if isinstance(node, ast.Constant)
        and isinstance(node.s, str)
        and node.s.startswith("/")
    ]

source = pathlib.Path("analysis.py").read_text()
print(find_hardcoded_paths(source))

Step 2: Refactor for Best Practices

Apply improvements in place. Always back up originals first.

2a. Add docstrings
python

# Before
def load_data(path):
    import pandas as pd
    return pd.read_csv(path)

# After
def load_data(path: str) -> "pd.DataFrame":
    """Load a CSV dataset from disk.

    Parameters
    ----------
    path : str
        Path to the CSV file (relative to project root).

    Returns
    -------
    pd.DataFrame
        Raw dataset with original column names preserved.
    """
    import pandas as pd
    return pd.read_csv(path)
2b. Parameterize hardcoded values
python
from pathlib import Path
import argparse

def parse_args():
    parser = argparse.ArgumentParser()
    parser.add_argument("--data", type=Path, default=Path("data/raw.csv"))
    parser.add_argument("--output", type=Path, default=Path("results/"))
    return parser.parse_args()

args = parse_args()
df = pd.read_csv(args.data)
args.output.mkdir(parents=True, exist_ok=True)
2c. Set random seeds
python
import random
import numpy as np

SEED = 42  # document this constant at module level

random.seed(SEED)
np.random.seed(SEED)

# scikit-learn
from sklearn.ensemble import RandomForestClassifier
clf = RandomForestClassifier(random_state=SEED)

# PyTorch
import torch
torch.manual_seed(SEED)
torch.backends.cudnn.deterministic = True
2d. Add error handling and logging
python
import logging
from pathlib import Path

logging.basicConfig(level=logging.INFO, format="%(asctime)s [%(levelname)s] %(message)s")
logger = logging.getLogger(__name__)

def load_data(path: Path) -> "pd.DataFrame":
    """Load dataset with validation."""
    import pandas as pd
    if not path.exists():
        raise FileNotFoundError(f"Data file not found: {path}")
    logger.info("Loading data from %s", path)
    df = pd.read_csv(path)
    if df.empty:
        raise ValueError(f"Loaded dataframe is empty: {path}")
    logger.info("Loaded %d rows, %d columns", *df.shape)
    return df

Step 3: Generate Environment Specifications

See references/environment-setup.md for full Dockerfile and Conda environment templates.

Show full SKILL.md (384 more words)Show less
requirements.txt (pip)
text
pip install pipreqs
pipreqs src/ --output requirements.txt --force

Verify resolution:

text
python -m venv .venv_test && source .venv_test/bin/activate
pip install -r requirements.txt
python -c "import pandas, numpy, sklearn"
deactivate && rm -rf .venv_test
environment.yml (Conda)
yaml
name: my-research-env
channels:
  - conda-forge
  - defaults
dependencies:
  - python=3.9
  - numpy=1.24.3
  - pandas=2.0.1
  - scikit-learn=1.2.2
  - matplotlib=3.7.1
  - pip:
    - some-pip-only-package==0.5.0
text
conda env create -f environment.yml
conda activate my-research-env

Step 4: Create Documentation

README structure

Generate a README.md containing at minimum:

markdown

## Requirements
<!-- List Python version and key packages with versions -->

## Installation
```text
conda env create -f environment.yml
conda activate my-research-env

Data

<!-- Describe input data format, source, and where to place files -->

Running the Analysis

text
python main.py --data data/raw.csv --output results/

Expected Outputs

<!-- Describe files created and how to interpret them -->

Reproducing Results

  • Random seed: 42 (set in config.py)
  • Hardware: results validated on CPU; GPU results may differ slightly

---

## Step 5: Validate Reproducibility

After all changes, verify that behaviour is unchanged:

```text

# 1. Run the full pipeline and capture output checksums
python main.py --data data/raw.csv --output results/
md5sum results/*.csv > checksums_refactored.md5
diff checksums_original.md5 checksums_refactored.md5

# 2. Run unit tests
pytest tests/ -v --tb=short

# 3. Confirm determinism across two clean runs
python main.py --output results_run1/
python main.py --output results_run2/
diff -r results_run1/ results_run2/

Reproducibility verification checklist:

  • Output checksums match pre-refactor baseline
  • All tests pass
  • Pipeline runs twice and produces identical outputs
  • requirements.txt / environment.yml installs cleanly in a fresh environment
  • No absolute paths remain in source files
  • Random seeds are set and documented
  • All public functions have docstrings
  • README contains complete reproduction instructions

Best Practices Summary

Practice
Relative paths only
Pin dependency versions
Set random seeds
Docstrings on all public functions
Validate outputs against a baseline
Automate environment setup

References

  • references/guide.md — Comprehensive user guide
  • references/environment-setup.md — Dockerfile and full environment templates
  • references/examples/ — Working code examples
  • references/api-docs/ — Complete API documentation

Skill ID: 455 | Version: 1.0 | License: MIT

Output Requirements

Every final response should make these items explicit when they are relevant:

  • Objective or requested deliverable
  • Inputs used and assumptions introduced
  • Workflow or decision path
  • Core result, recommendation, or artifact
  • Constraints, risks, caveats, or validation needs
  • Unresolved items and next-step checks

Error Handling

  • If required inputs are missing, state exactly which fields are missing and request only the minimum additional information.
  • If the task goes outside the documented scope, stop instead of guessing or silently widening the assignment.
  • If scripts/main.py fails, report the failure point, summarize what still can be completed safely, and provide a manual fallback.
  • Do not fabricate files, citations, data, search results, or execution outcomes.

Input Validation

This skill accepts requests that match the documented purpose of code-refactor-for-reproducibility and include enough context to complete the workflow safely.

Do not continue the workflow when the request is out of scope, missing a critical input, or would require unsupported assumptions. Instead respond:

code-refactor-for-reproducibility only handles its documented workflow. Please provide the missing required inputs or switch to a more suitable skill.

Response Template

Use the following fixed structure for non-trivial requests:

  1. Objective
  2. Inputs Received
  3. Assumptions
  4. Workflow
  5. Deliverable
  6. Risks and Limits
  7. Next Checks

If the request is simple, you may compress the structure, but still keep assumptions and limits explicit when they affect correctness.

© aipoch, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (scripts) in scientific-skills/Data Analysis/code-refactor-for-reproducibility of aipoch/medical-research-skills.

  • SKILL.md
  • code-refactor-for-reproducibility_audit_result_v2.json
  • requirements.txt
  • scripts/main.py
  • tile.json

Open the folder on GitHubat commit 686e09d

Compare with similar skills

Code Refactor For Reproducibility next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Code Refactor For Reproducibility compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Code Refactor For Reproducibility this skillaipoch/medical-research-skills2k—~3.1kAutomated safety check: PassMIT
Writing Lean Proofstrailofbits/skills7.4k—~4kAutomated safety check: PassCC-BY-SA-4.0
Setting Up Reproducible AnalysisK-Dense-AI/science-superpowers348—~1.5kAutomated safety check: PassCustom licence
Review Rpedrohcgs/claude-code-my-workflow1.6k—~430Automated safety check: PassMIT
Bio Workflow Management Wdl WorkflowsGPTomics/bioSkills1.2k1 repos~3.5kAutomated safety check: PassMIT
Review Rbrycewang-stanford/Auto-Empirical-Research-Skills4.5k—~1.3kAutomated safety check: PassCustom licence

Similar skills

  • Writing Lean Proofs

    trailofbits/skills

    Official

    Structures Lean 4 proofs and library design along Mathlib conventions, from stating theorems to refactoring long tactic proofs and fixing slow or timing-out ones.

    7.4k GitHub stars~4k tokensUpdated yesterday
    DevelopmentAuto-check passed
  • Setting Up Reproducible Analysis

    K-Dense-AI/science-superpowers

    A skill your agent uses when starting analysis work that needs isolation, or before executing a pre-registered plan - ensures an isolated, reproducible workspace with pinned environment, fixed…

    348 GitHub stars~1.5k tokensUpdated 25 days ago
    DevelopmentAuto-check passed
  • Review R

    pedrohcgs/claude-code-my-workflow

    Read-only R code review protocol for .R scripts. An agent skill from pedrohcgs/claude-code-my-workflow.

    1.6k GitHub stars~430 tokensUpdated 11 days ago
    DevelopmentAuto-check passed
  • Authors bioinformatics pipelines in WDL (Workflow Description Language) run by Cromwell or miniwdl, targeting the GATK/Broad and Terra/AnVIL/BioData Catalyst cloud ecosystem, with tasks, workflows…

    1.2k GitHub starsUsed in 1 repo~3.5k tokens
    DevelopmentAuto-check passed
  • Review R

    brycewang-stanford/Auto-Empirical-Research-Skills

    R code review for the sewage project. An agent skill from brycewang-stanford/Auto-Empirical-Research-Skills.

    4.5k GitHub stars~1.3k tokensUpdated 3 days ago
    DevelopmentAuto-check passed
  • Guidelines

    akash-network/node

    Behavioral guidelines to reduce common LLM coding mistakes. An agent skill from akash-network/node.

    1.1k GitHub starsUsed in 22 repos~577 tokens
    DevelopmentAuto-check passed

More from aipoch/medical-research-skills

All 567 skills in this repo
  • Academic Poster Generator

    aipoch/medical-research-skills

    Complete workflow for generating academic research posters from PDF literature; use when you need to extract paper content from PDFs and produce a LaTeX-based poster…

    2k GitHub stars~2.2k tokensUpdated 21 days ago
    Auto-check passed
  • Diagnostic Study Quality Assessment Quadas

    aipoch/medical-research-skills

    Analyzes clinical diagnostic accuracy studies for bias using the QUADAS-2 tool.

    2k GitHub stars~1.4k tokensUpdated 21 days ago
    Auto-check passed
  • Exploratory Data Analysis

    aipoch/medical-research-skills

    Perform comprehensive exploratory data analysis on scientific data files across 200+ file formats.

    2k GitHub stars~3.7k tokensUpdated 21 days ago
    Auto-check passed
  • Iso Certification

    aipoch/medical-research-skills

    A toolkit for preparing ISO 13485:2016 certification documentation for medical device QMS.

    2k GitHub stars~1.8k tokensUpdated 21 days ago
    Auto-check passed
  • Journal Skills

    aipoch/medical-research-skills

    Recommends target journals for manuscript submission by analyzing the paper topic/abstract and the journal distribution of similar PubMed literature; use when users ask for journal…

    2k GitHub stars~1.7k tokensUpdated 21 days ago
    Auto-check passed
  • Latex Posters

    aipoch/medical-research-skills

    Creates academic-poster writing packages for LaTeX using beamerposter, tikzposter, or baposter.

    2k GitHub stars~1.3k tokensUpdated 21 days ago
    Auto-check passed

Questions about Code Refactor For Reproducibility

What does Code Refactor For Reproducibility do?

A skill your agent uses when refactoring research code for publication, adding documentation to existing analysis scripts, creating reproducible computational workflows, or preparing code for…. Code Refactor For Reproducibility is an agent skill from aipoch/medical-research-skills. Use when refactoring research code for publication, adding documentation to existing analysis scripts, creating reproducible computational workflows, or preparing code for sharing with collaborators.

When should I use Code Refactor For Reproducibility?

Code Refactor For Reproducibility fits situations like: refactoring research code for publication; adding documentation to existing analysis scripts; creating reproducible computational workflows; preparing code for sharing with collaborators.

How do I install Code Refactor For Reproducibility in Claude Code?

Run `npx skills add aipoch/medical-research-skills --skill code-refactor-for-reproducibility -a claude-code`. Or copy the skill folder (scientific-skills/Data Analysis/code-refactor-for-reproducibility in aipoch/medical-research-skills) into .claude/skills/code-refactor-for-reproducibility in your project. Claude Code loads it when a task matches its description.

How do I install Code Refactor For Reproducibility in Codex?

Run `npx skills add aipoch/medical-research-skills --skill code-refactor-for-reproducibility -a codex`. Or copy the skill folder (scientific-skills/Data Analysis/code-refactor-for-reproducibility in aipoch/medical-research-skills) into .agents/skills/code-refactor-for-reproducibility in your project. Codex loads it when a task matches its description.

Can I use Code Refactor For Reproducibility in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add aipoch/medical-research-skills --skill code-refactor-for-reproducibility -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/code-refactor-for-reproducibility, .gemini/skills/code-refactor-for-reproducibility, .github/skills/code-refactor-for-reproducibility and .opencode/skills/code-refactor-for-reproducibility in your project.

What does Code Refactor For Reproducibility need to run?

Going by SKILL.md and its folder, Code Refactor For Reproducibility needs Python for the scripts in its folder and the command-line tools its instructions call (python and pytest). Our summary lists: Python 3.

Does Code Refactor For Reproducibility access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Code Refactor For Reproducibility safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Code Refactor For Reproducibility use?

Code Refactor For Reproducibility is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Code Refactor For Reproducibility use?

About 3.1k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Code Refactor For Reproducibility?

Skills that share tags, products or a category with Code Refactor For Reproducibility: Writing Lean Proofs (trailofbits/skills, 7.4k stars), Setting Up Reproducible Analysis (K-Dense-AI/science-superpowers, 348 stars), Review R (pedrohcgs/claude-code-my-workflow, 1.6k stars) and Bio Workflow Management Wdl Workflows (GPTomics/bioSkills, 1.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Code Refactor For Reproducibility?

aipoch (a GitHub organization) maintains it in aipoch/medical-research-skills, which has 1,974 GitHub stars. The repository holds 567 skills in this directory. The repository was last updated on September 17, 2026.

Source: aipoch/medical-research-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.