Reproducible Pipelines
brycewang-stanford/Auto-Empirical-Research-Skills
This skill covers reproducible research pipelines and replication packages.
Build reproducible data science pipelines with Kedro for research projects
$ npx skills add wentorai/research-plugins --skill kedro-pipeline-guide -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install wentorai/research-plugins kedro-pipeline-guide --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/wentorai/research-plugins.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/research/automation/kedro-pipeline-guide .claude/skills/kedro-pipeline-guide && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "kedro-pipeline-guide" agent skill from https://github.com/wentorai/research-plugins/tree/main/skills/research/automation/kedro-pipeline-guide into .claude/skills/kedro-pipeline-guide/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "kedro-pipeline-guide", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/wentorai/research-plugins/tree/main/skills/research/automation/kedro-pipeline-guideType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add wentorai/research-plugins --skill kedro-pipeline-guide -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install wentorai/research-plugins kedro-pipeline-guide --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/wentorai/research-plugins.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/research/automation/kedro-pipeline-guide .agents/skills/kedro-pipeline-guide && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "kedro-pipeline-guide" agent skill from https://github.com/wentorai/research-plugins/tree/main/skills/research/automation/kedro-pipeline-guide into .agents/skills/kedro-pipeline-guide/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "kedro-pipeline-guide", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add wentorai/research-plugins --skill kedro-pipeline-guide -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install wentorai/research-plugins kedro-pipeline-guide --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/wentorai/research-plugins.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/research/automation/kedro-pipeline-guide .cursor/skills/kedro-pipeline-guide && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "kedro-pipeline-guide" agent skill from https://github.com/wentorai/research-plugins/tree/main/skills/research/automation/kedro-pipeline-guide into .cursor/skills/kedro-pipeline-guide/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "kedro-pipeline-guide", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/wentorai/research-plugins.git --path skills/research/automation/kedro-pipeline-guide--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add wentorai/research-plugins --skill kedro-pipeline-guide -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install wentorai/research-plugins kedro-pipeline-guide --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/wentorai/research-plugins.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/research/automation/kedro-pipeline-guide .gemini/skills/kedro-pipeline-guide && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "kedro-pipeline-guide" agent skill from https://github.com/wentorai/research-plugins/tree/main/skills/research/automation/kedro-pipeline-guide into .gemini/skills/kedro-pipeline-guide/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "kedro-pipeline-guide", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install wentorai/research-plugins kedro-pipeline-guideInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add wentorai/research-plugins --skill kedro-pipeline-guide -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/wentorai/research-plugins.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/research/automation/kedro-pipeline-guide .github/skills/kedro-pipeline-guide && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "kedro-pipeline-guide" agent skill from https://github.com/wentorai/research-plugins/tree/main/skills/research/automation/kedro-pipeline-guide into .github/skills/kedro-pipeline-guide/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "kedro-pipeline-guide", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add wentorai/research-plugins --skill kedro-pipeline-guide -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install wentorai/research-plugins kedro-pipeline-guide --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/wentorai/research-plugins.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/research/automation/kedro-pipeline-guide .opencode/skills/kedro-pipeline-guide && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "kedro-pipeline-guide" agent skill from https://github.com/wentorai/research-plugins/tree/main/skills/research/automation/kedro-pipeline-guide into .opencode/skills/kedro-pipeline-guide/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "kedro-pipeline-guide", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
kedro-pipeline-guideBuild reproducible data science pipelines with Kedro for research projects
Kedro Pipeline Guide is an agent skill from wentorai/research-plugins. Build reproducible data science pipelines with Kedro for research projects
Its SKILL.md is about 1.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
The repository describes itself as: 350+ academic research skills, MCP configs, and plugins for Research-Claw and AI agents. The licence is MIT.
Read from SKILL.md and the folder at commit bf44b3c. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pipFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
github.comdocs.kedro.orgFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Kedro Pipeline Guide loads about 1.8k tokens when it runs. Until then it costs about 24 tokens; SKILL.md has 436 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from wentorai/research-plugins at commit bf44b3c, republished under its MIT licence (© wentorai). 436 words, ~1,759 tokens.
.claude/skills/kedro-pipeline-guide/SKILL.md (or your agent's skills folder).Kedro is an open-source Python framework for creating reproducible, maintainable, and modular data science pipelines. Developed originally at McKinsey's QuantumBlack labs, Kedro provides an opinionated project structure and a set of conventions that transform ad-hoc analysis scripts into production-quality code that can be tested, versioned, and shared across research teams.
In academic research, reproducibility is both a scientific imperative and a practical challenge. Jupyter notebooks and standalone scripts often become tangled webs of dependencies that are difficult to re-run months later when responding to reviewer comments or extending prior work. Kedro addresses this by separating data processing logic from data access, enforcing explicit pipeline definitions, and providing built-in data versioning and experiment tracking.
With over 11,000 GitHub stars, Kedro has gained adoption across industry and academia. Its design philosophy aligns naturally with the needs of computational research: clear data lineage, parameterized experiments, and the ability to scale from a laptop to a cluster without rewriting code.
Install Kedro via pip:
pip install kedroCreate a new project using the Kedro starter:
kedro new --name my-research-project --tools lint,test,docs
cd my-research-projectThis generates a standardized project structure:
my-research-project/
conf/
base/
catalog.yml # Data source definitions
parameters.yml # Experiment parameters
local/ # Local overrides (gitignored)
src/
my_research_project/
pipelines/
data_processing/
nodes.py # Pure Python functions
pipeline.py # Pipeline definition
modeling/
nodes.py
pipeline.py
pipeline_registry.py
data/ # Local data directory
notebooks/ # Jupyter notebooks
tests/ # Unit testsInstall project dependencies:
pip install -e ".[dev]"Nodes: The fundamental units of computation in Kedro. Each node is a pure Python function with explicitly declared inputs and outputs:
# src/my_research_project/pipelines/data_processing/nodes.py
import pandas as pd
from sklearn.preprocessing import StandardScaler
def clean_raw_data(raw_data: pd.DataFrame) -> pd.DataFrame:
"""Remove missing values and outliers from raw experimental data."""
cleaned = raw_data.dropna(subset=["measurement", "condition"])
q1 = cleaned["measurement"].quantile(0.01)
q99 = cleaned["measurement"].quantile(0.99)
return cleaned[cleaned["measurement"].between(q1, q99)]
def normalize_features(
cleaned_data: pd.DataFrame, parameters: dict
) -> pd.DataFrame:
"""Standardize feature columns specified in parameters."""
feature_cols = parameters["feature_columns"]
scaler = StandardScaler()
result = cleaned_data.copy()
result[feature_cols] = scaler.fit_transform(cleaned_data[feature_cols])
return resultPipelines: Chains of nodes connected through named datasets:
# src/my_research_project/pipelines/data_processing/pipeline.py
from kedro.pipeline import Pipeline, node, pipeline
from .nodes import clean_raw_data, normalize_features
def create_pipeline(**kwargs) -> Pipeline:
return pipeline([
node(
func=clean_raw_data,
inputs="raw_experiment_data",
outputs="cleaned_data",
name="clean_data_node",
),
node(
func=normalize_features,
inputs=["cleaned_data", "params:preprocessing"],
outputs="normalized_data",
name="normalize_node",
),
])Data Catalog: A declarative registry that maps logical dataset names to physical storage:
# conf/base/catalog.yml
raw_experiment_data:
type: pandas.CSVDataset
filepath: data/01_raw/experiment_results.csv
cleaned_data:
type: pandas.ParquetDataset
filepath: data/02_intermediate/cleaned.parquet
normalized_data:
type: pandas.ParquetDataset
filepath: data/03_primary/normalized.parquet
versioned: trueThe versioned: true flag automatically creates timestamped versions of outputs, enabling exact reproduction of prior runs.
Parameters: Experiment configuration separated from code:
# conf/base/parameters.yml
preprocessing:
feature_columns:
- temperature
- pressure
- concentration
outlier_method: iqr
modeling:
algorithm: random_forest
n_estimators: 500
max_depth: 10
test_size: 0.2
random_seed: 42Execute the full pipeline:
kedro runRun a specific pipeline or node:
kedro run --pipeline data_processing
kedro run --nodes clean_data_nodeVisualize the pipeline dependency graph:
pip install kedro-viz
kedro viz runThis launches an interactive web visualization showing the complete data flow, making it easy to understand and communicate your analytical pipeline to collaborators and reviewers.
Experiment Reproducibility: Every Kedro run uses explicit parameters and versioned data. Store parameter files in Git alongside code to create a complete record of every experiment configuration.
Reviewer Response: When peer reviewers request additional analyses or modified parameters, change parameters.yml and re-run. The pipeline automatically reprocesses only affected downstream nodes.
Team Collaboration: Multiple researchers can work on different pipeline modules simultaneously. The explicit input/output contracts between nodes prevent integration conflicts.
Scaling Computation: Kedro pipelines can be deployed to distributed computing platforms without code changes using runners:
# Run with parallel execution
kedro run --runner=ParallelRunner
# Deploy to Airflow, Prefect, or other orchestrators
pip install kedro-airflow
kedro airflow createIntegration with Jupyter: Use Kedro notebooks for exploration while maintaining the pipeline for production runs:
kedro jupyter notebookThe Kedro Jupyter integration automatically loads the project catalog and parameters, bridging the gap between interactive exploration and pipeline execution.
© wentorai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/research/automation/kedro-pipeline-guide of wentorai/research-plugins.
Open the folder on GitHubat commit bf44b3c
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in wentorai/research-plugins, which our catalogue first saw on October 7, 2026.
Kedro Pipeline Guide next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Kedro Pipeline Guide this skillwentorai/research-plugins | 298 | 1 repos | ~1.8k | Automated safety check: Pass | MIT | |
| Reproducible Pipelinesbrycewang-stanford/Auto-Empirical-Research-Skills | 4.6k | — | ~3.3k | Automated safety check: Pass | Custom licence | |
| Opensource Pipelineaffaan-m/ECC | 277k | 1 repos | ~2.2k | Automated safety check: Notes | MIT | |
| Orch Pipelineaffaan-m/ECC | 276k | 1 repos | ~1.6k | Automated safety check: Pass | MIT | |
| AI Pipeline Orchestrationsickn33/agentic-awesome-skills | 47k | 1 repos | ~2.5k | Automated safety check: Pass | MIT | |
| Data PipelineRightNow-AI/openfang | 18k | — | ~847 | Automated safety check: Pass | Apache-2.0 |
brycewang-stanford/Auto-Empirical-Research-Skills
This skill covers reproducible research pipelines and replication packages.
affaan-m/ECC
Open-source pipeline: fork, sanitize, and package private projects for safe public release.
affaan-m/ECC
Shared orchestration engine behind the orch- skill family — the gated Research-Plan-TDD-Review-Commit pipeline, size classifier, agent and command map, and two human gates (plan approval, commit…
sickn33/agentic-awesome-skills
Orchestrate AI/ML pipelines for data ingestion, model training, batch inference, and RAG indexing using Prefect, Airflow, or Dagster.
RightNow-AI/openfang
Data pipeline expert for ETL, Apache Spark, Airflow, dbt, and data quality
windmill-labs/windmill
Builds Windmill data pipelines as independent, pipeline-annotated scripts forming a DAG over shared storage, defaulting to DuckDB nodes that materialize into DuckLake tables.
wentorai/research-plugins
Craft structured research abstracts that maximize clarity and journal acceptance
wentorai/research-plugins
Manage academic citations across BibTeX, APA, MLA, and Chicago formats
wentorai/research-plugins
Summarize academic papers with structured extraction of key elements
wentorai/research-plugins
Evidence-based study techniques for academic learning and retention
wentorai/research-plugins
Adjust writing tone and register for academic audiences and venues
wentorai/research-plugins
Academic translation, post-editing, and Chinglish correction guide
Build reproducible data science pipelines with Kedro for research projects. Kedro Pipeline Guide is an agent skill from wentorai/research-plugins.
Run `npx skills add wentorai/research-plugins --skill kedro-pipeline-guide -a claude-code`. Or copy the skill folder (skills/research/automation/kedro-pipeline-guide in wentorai/research-plugins) into .claude/skills/kedro-pipeline-guide in your project. Claude Code loads it when a task matches its description.
Run `npx skills add wentorai/research-plugins --skill kedro-pipeline-guide -a codex`. Or copy the skill folder (skills/research/automation/kedro-pipeline-guide in wentorai/research-plugins) into .agents/skills/kedro-pipeline-guide in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add wentorai/research-plugins --skill kedro-pipeline-guide -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/kedro-pipeline-guide, .gemini/skills/kedro-pipeline-guide, .github/skills/kedro-pipeline-guide and .opencode/skills/kedro-pipeline-guide in your project.
Going by SKILL.md and its folder, Kedro Pipeline Guide needs the command-line tools its instructions call (pip). Our summary lists: Python 3.
SKILL.md names 2 domains. As links in the text: github.com and docs.kedro.org. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Kedro Pipeline Guide is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.8k tokens (SKILL.md is roughly 7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Kedro Pipeline Guide: Reproducible Pipelines (brycewang-stanford/Auto-Empirical-Research-Skills, 4.6k stars), Opensource Pipeline (affaan-m/ECC, 277k stars), Orch Pipeline (affaan-m/ECC, 276k stars) and AI Pipeline Orchestration (sickn33/agentic-awesome-skills, 47k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
wentorai (a GitHub user) maintains it in wentorai/research-plugins, which has 298 GitHub stars. The repository holds 405 skills in this directory. The repository was last updated on June 19, 2026.
Source: wentorai/research-plugins on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.