Agent skill

Kedro Pipeline Guide

by wentorai in wentorai/research-plugins

Build reproducible data science pipelines with Kedro for research projects

MITAuto-check passed

Install Kedro Pipeline Guide

skills CLI
$ npx skills add wentorai/research-plugins --skill kedro-pipeline-guide -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install wentorai/research-plugins kedro-pipeline-guide --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/wentorai/research-plugins.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/research/automation/kedro-pipeline-guide .claude/skills/kedro-pipeline-guide && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
kedro-pipeline-guide
GitHub stars
298
Used in
1 other repo
Token cost
~1.8k tokens
SKILL.md length
436 words
Files
1
Skills in repo
405
Repo updated
First seen
Licence
MIT

At a glance

Build reproducible data science pipelines with Kedro for research projects

  • SKILL.md covers Overview, Installation and Setup, Core Concepts and Running and Visualizing…, plus 2 more sections
  • Calls pip

What it does

Kedro Pipeline Guide is an agent skill from wentorai/research-plugins. Build reproducible data science pipelines with Kedro for research projects

Its SKILL.md is about 1.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

The repository describes itself as: 350+ academic research skills, MCP configs, and plugins for Research-Claw and AI agents. The licence is MIT.

Example prompts

  • “/kedro-pipeline-guide”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit bf44b3c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com
    • docs.kedro.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Kedro Pipeline Guide loads about 1.8k tokens when it runs. Until then it costs about 24 tokens; SKILL.md has 436 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~24
When it runs · the whole SKILL.md, loaded when a task matches
~1.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from wentorai/research-plugins at commit bf44b3c, republished under its MIT licence (© wentorai). 436 words, ~1,759 tokens.

Download SKILL.mdSave it as .claude/skills/kedro-pipeline-guide/SKILL.md (or your agent's skills folder).
name
kedro-pipeline-guide
description
Build reproducible data science pipelines with Kedro for research projects

Kedro Pipeline Guide

Overview

Kedro is an open-source Python framework for creating reproducible, maintainable, and modular data science pipelines. Developed originally at McKinsey's QuantumBlack labs, Kedro provides an opinionated project structure and a set of conventions that transform ad-hoc analysis scripts into production-quality code that can be tested, versioned, and shared across research teams.

In academic research, reproducibility is both a scientific imperative and a practical challenge. Jupyter notebooks and standalone scripts often become tangled webs of dependencies that are difficult to re-run months later when responding to reviewer comments or extending prior work. Kedro addresses this by separating data processing logic from data access, enforcing explicit pipeline definitions, and providing built-in data versioning and experiment tracking.

With over 11,000 GitHub stars, Kedro has gained adoption across industry and academia. Its design philosophy aligns naturally with the needs of computational research: clear data lineage, parameterized experiments, and the ability to scale from a laptop to a cluster without rewriting code.

Installation and Setup

Install Kedro via pip:

bash
pip install kedro

Create a new project using the Kedro starter:

bash
kedro new --name my-research-project --tools lint,test,docs
cd my-research-project

This generates a standardized project structure:

my-research-project/
  conf/
    base/
      catalog.yml        # Data source definitions
      parameters.yml     # Experiment parameters
    local/               # Local overrides (gitignored)
  src/
    my_research_project/
      pipelines/
        data_processing/
          nodes.py       # Pure Python functions
          pipeline.py    # Pipeline definition
        modeling/
          nodes.py
          pipeline.py
      pipeline_registry.py
  data/                  # Local data directory
  notebooks/             # Jupyter notebooks
  tests/                 # Unit tests

Install project dependencies:

bash
pip install -e ".[dev]"

Core Concepts

Nodes: The fundamental units of computation in Kedro. Each node is a pure Python function with explicitly declared inputs and outputs:

python
# src/my_research_project/pipelines/data_processing/nodes.py

import pandas as pd
from sklearn.preprocessing import StandardScaler

def clean_raw_data(raw_data: pd.DataFrame) -> pd.DataFrame:
    """Remove missing values and outliers from raw experimental data."""
    cleaned = raw_data.dropna(subset=["measurement", "condition"])
    q1 = cleaned["measurement"].quantile(0.01)
    q99 = cleaned["measurement"].quantile(0.99)
    return cleaned[cleaned["measurement"].between(q1, q99)]

def normalize_features(
    cleaned_data: pd.DataFrame, parameters: dict
) -> pd.DataFrame:
    """Standardize feature columns specified in parameters."""
    feature_cols = parameters["feature_columns"]
    scaler = StandardScaler()
    result = cleaned_data.copy()
    result[feature_cols] = scaler.fit_transform(cleaned_data[feature_cols])
    return result

Pipelines: Chains of nodes connected through named datasets:

python
# src/my_research_project/pipelines/data_processing/pipeline.py

from kedro.pipeline import Pipeline, node, pipeline
from .nodes import clean_raw_data, normalize_features

def create_pipeline(**kwargs) -> Pipeline:
    return pipeline([
        node(
            func=clean_raw_data,
            inputs="raw_experiment_data",
            outputs="cleaned_data",
            name="clean_data_node",
        ),
        node(
            func=normalize_features,
            inputs=["cleaned_data", "params:preprocessing"],
            outputs="normalized_data",
            name="normalize_node",
        ),
    ])

Data Catalog: A declarative registry that maps logical dataset names to physical storage:

yaml
# conf/base/catalog.yml
raw_experiment_data:
  type: pandas.CSVDataset
  filepath: data/01_raw/experiment_results.csv

cleaned_data:
  type: pandas.ParquetDataset
  filepath: data/02_intermediate/cleaned.parquet

normalized_data:
  type: pandas.ParquetDataset
  filepath: data/03_primary/normalized.parquet
  versioned: true

The versioned: true flag automatically creates timestamped versions of outputs, enabling exact reproduction of prior runs.

Parameters: Experiment configuration separated from code:

yaml
# conf/base/parameters.yml
preprocessing:
  feature_columns:
    - temperature
    - pressure
    - concentration
  outlier_method: iqr

modeling:
  algorithm: random_forest
  n_estimators: 500
  max_depth: 10
  test_size: 0.2
  random_seed: 42
Show full SKILL.md (186 more words)Show less

Running and Visualizing Pipelines

Execute the full pipeline:

bash
kedro run

Run a specific pipeline or node:

bash
kedro run --pipeline data_processing
kedro run --nodes clean_data_node

Visualize the pipeline dependency graph:

bash
pip install kedro-viz
kedro viz run

This launches an interactive web visualization showing the complete data flow, making it easy to understand and communicate your analytical pipeline to collaborators and reviewers.

Research Workflow Integration

Experiment Reproducibility: Every Kedro run uses explicit parameters and versioned data. Store parameter files in Git alongside code to create a complete record of every experiment configuration.

Reviewer Response: When peer reviewers request additional analyses or modified parameters, change parameters.yml and re-run. The pipeline automatically reprocesses only affected downstream nodes.

Team Collaboration: Multiple researchers can work on different pipeline modules simultaneously. The explicit input/output contracts between nodes prevent integration conflicts.

Scaling Computation: Kedro pipelines can be deployed to distributed computing platforms without code changes using runners:

bash
# Run with parallel execution
kedro run --runner=ParallelRunner

# Deploy to Airflow, Prefect, or other orchestrators
pip install kedro-airflow
kedro airflow create

Integration with Jupyter: Use Kedro notebooks for exploration while maintaining the pipeline for production runs:

bash
kedro jupyter notebook

The Kedro Jupyter integration automatically loads the project catalog and parameters, bridging the gap between interactive exploration and pipeline execution.

References

© wentorai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/research/automation/kedro-pipeline-guide of wentorai/research-plugins.

Open the folder on GitHubat commit bf44b3c

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in wentorai/research-plugins, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Kedro Pipeline Guide next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Kedro Pipeline Guide compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Kedro Pipeline Guide this skillwentorai/research-plugins2981 repos~1.8kAutomated safety check: PassMIT
Reproducible Pipelinesbrycewang-stanford/Auto-Empirical-Research-Skills4.6k—~3.3kAutomated safety check: PassCustom licence
Opensource Pipelineaffaan-m/ECC277k1 repos~2.2kAutomated safety check: NotesMIT
Orch Pipelineaffaan-m/ECC276k1 repos~1.6kAutomated safety check: PassMIT
AI Pipeline Orchestrationsickn33/agentic-awesome-skills47k1 repos~2.5kAutomated safety check: PassMIT
Data PipelineRightNow-AI/openfang18k—~847Automated safety check: PassApache-2.0

Similar skills

  • Reproducible Pipelines

    brycewang-stanford/Auto-Empirical-Research-Skills

    This skill covers reproducible research pipelines and replication packages.

    4.6k GitHub stars~3.3k tokensUpdated 5 days ago
    Research & ScienceAuto-check passed
  • Open-source pipeline: fork, sanitize, and package private projects for safe public release.

    277k GitHub starsUsed in 1 repo~2.2k tokens
    Auto-check: notes
  • Orch Pipeline

    affaan-m/ECC

    Shared orchestration engine behind the orch- skill family — the gated Research-Plan-TDD-Review-Commit pipeline, size classifier, agent and command map, and two human gates (plan approval, commit…

    276k GitHub starsUsed in 1 repo~1.6k tokens
    Testing & QAAuto-check passed
  • AI Pipeline Orchestration

    sickn33/agentic-awesome-skills

    Orchestrate AI/ML pipelines for data ingestion, model training, batch inference, and RAG indexing using Prefect, Airflow, or Dagster.

    47k GitHub starsUsed in 1 repo~2.5k tokens
    Data & AnalyticsAuto-check passed
  • Data Pipeline

    RightNow-AI/openfang

    Data pipeline expert for ETL, Apache Spark, Airflow, dbt, and data quality

    18k GitHub stars~847 tokensUpdated 3 mo ago
    Data & AnalyticsAuto-check passed
  • Windmill Data Pipeline Author

    windmill-labs/windmill

    Builds Windmill data pipelines as independent, pipeline-annotated scripts forming a DAG over shared storage, defaulting to DuckDB nodes that materialize into DuckLake tables.

    18k GitHub stars~4k tokensUpdated today
    Data & AnalyticsAuto-check passed

More from wentorai/research-plugins

All 405 skills in this repo
  • Abstract Writing Guide

    wentorai/research-plugins

    Craft structured research abstracts that maximize clarity and journal acceptance

    298 GitHub starsUsed in 1 repo~1.7k tokens
    Auto-check passed
  • Academic Citation Manager

    wentorai/research-plugins

    Manage academic citations across BibTeX, APA, MLA, and Chicago formats

    298 GitHub starsUsed in 1 repo~2.7k tokens
    Auto-check passed
  • Academic Paper Summarizer

    wentorai/research-plugins

    Summarize academic papers with structured extraction of key elements

    298 GitHub starsUsed in 1 repo~1.4k tokens
    Auto-check passed
  • Academic Study Methods

    wentorai/research-plugins

    Evidence-based study techniques for academic learning and retention

    298 GitHub starsUsed in 1 repo~1.8k tokens
    Auto-check passed
  • Academic Tone Guide

    wentorai/research-plugins

    Adjust writing tone and register for academic audiences and venues

    298 GitHub starsUsed in 1 repo~1.9k tokens
    Auto-check passed
  • Academic Translation Guide

    wentorai/research-plugins

    Academic translation, post-editing, and Chinglish correction guide

    298 GitHub starsUsed in 1 repo~1.6k tokens
    Auto-check passed

Questions about Kedro Pipeline Guide

What does Kedro Pipeline Guide do?

Build reproducible data science pipelines with Kedro for research projects. Kedro Pipeline Guide is an agent skill from wentorai/research-plugins.

How do I install Kedro Pipeline Guide in Claude Code?

Run `npx skills add wentorai/research-plugins --skill kedro-pipeline-guide -a claude-code`. Or copy the skill folder (skills/research/automation/kedro-pipeline-guide in wentorai/research-plugins) into .claude/skills/kedro-pipeline-guide in your project. Claude Code loads it when a task matches its description.

How do I install Kedro Pipeline Guide in Codex?

Run `npx skills add wentorai/research-plugins --skill kedro-pipeline-guide -a codex`. Or copy the skill folder (skills/research/automation/kedro-pipeline-guide in wentorai/research-plugins) into .agents/skills/kedro-pipeline-guide in your project. Codex loads it when a task matches its description.

Can I use Kedro Pipeline Guide in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add wentorai/research-plugins --skill kedro-pipeline-guide -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/kedro-pipeline-guide, .gemini/skills/kedro-pipeline-guide, .github/skills/kedro-pipeline-guide and .opencode/skills/kedro-pipeline-guide in your project.

What does Kedro Pipeline Guide need to run?

Going by SKILL.md and its folder, Kedro Pipeline Guide needs the command-line tools its instructions call (pip). Our summary lists: Python 3.

Does Kedro Pipeline Guide access the network?

SKILL.md names 2 domains. As links in the text: github.com and docs.kedro.org. This is read from the text; nothing was executed.

Is Kedro Pipeline Guide safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Kedro Pipeline Guide use?

Kedro Pipeline Guide is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Kedro Pipeline Guide use?

About 1.8k tokens (SKILL.md is roughly 7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Kedro Pipeline Guide?

Skills that share tags, products or a category with Kedro Pipeline Guide: Reproducible Pipelines (brycewang-stanford/Auto-Empirical-Research-Skills, 4.6k stars), Opensource Pipeline (affaan-m/ECC, 277k stars), Orch Pipeline (affaan-m/ECC, 276k stars) and AI Pipeline Orchestration (sickn33/agentic-awesome-skills, 47k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Kedro Pipeline Guide?

wentorai (a GitHub user) maintains it in wentorai/research-plugins, which has 298 GitHub stars. The repository holds 405 skills in this directory. The repository was last updated on June 19, 2026.

Source: wentorai/research-plugins on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.