Agent skill

Preprocessing Data With Automated Pipelines

by foryourhealth111-pixel in foryourhealth111-pixel/Vibe-Skills

Design and implement repeatable preprocessing pipelines for cleaning, encoding, transforming, and validating ML input data.

MITAuto-check passedData & Analytics

Install Preprocessing Data With Automated Pipelines

skills CLI
$ npx skills add foryourhealth111-pixel/Vibe-Skills --skill preprocessing-data-with-automated-pipelines -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install foryourhealth111-pixel/Vibe-Skills preprocessing-data-with-automated-pipelines --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/foryourhealth111-pixel/Vibe-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/bundled/skills/preprocessing-data-with-automated-pipelines .claude/skills/preprocessing-data-with-automated-pipelines && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
preprocessing-data-with-automated-pipelines
GitHub stars
3.6k
Token cost
~374 tokens
SKILL.md length
136 words
Files
9 (incl. scripts, references, assets)
Skills in repo
81
Repo updated
First seen
Licence
MIT

At a glance

Design and implement repeatable preprocessing pipelines for cleaning, encoding, transforming, and validating ML input data.

  • Tasks that involve Machine learning
  • SKILL.md covers Positioning, When to Use, Not For / Boundaries and Typical Outputs, plus 1 more section
  • Runs Python scripts from its folder

What it does

Preprocessing Data With Automated Pipelines is an agent skill from foryourhealth111-pixel/Vibe-Skills. Design and implement repeatable preprocessing pipelines for cleaning, encoding, transforming, and validating ML input data.

Its SKILL.md is about 370 tokens, which your agent loads only when the skill is triggered. The skill folder holds 11 other files, including scripts, reference files and assets (for example `assets/README.md`, `references/README.md` and `scripts/README.md`).

It sits in Data & Analytics, covering Machine learning. The repository describes itself as: Intelligent Skill routing and workflow orchestration for AI agents — +21.12 pp reward, −29.6% tokens on SkillsBench with DeepSeekV4Flash-VE. The licence is MIT.

When your agent uses it

  • Tasks that involve Machine learning

Example prompts

  • “/preprocessing-data-with-automated-pipelines”

Requirements

  • Python 3
  • Pre-approved tools (allowed-tools): Read, Write, Edit, Grep, Glob, Bash(cmd:*)

What it can do on your machine

Read from SKILL.md and the folder at commit ddcaa2a. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Write
    • Edit
    • Grep
    • Glob
    • Bash(cmd:*)

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 5 files in scripts/ (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Preprocessing Data With Automated Pipelines loads about 374 tokens when it runs, and up to ~484 if it reads all its reference files. Until then it costs about 42 tokens; SKILL.md has 136 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~42
When it runs · the whole SKILL.md, loaded when a task matches
~374
With references · SKILL.md plus every file in references/, read only if the agent opens them
~484

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from foryourhealth111-pixel/Vibe-Skills at commit ddcaa2a, republished under its MIT licence (© foryourhealth111-pixel). 136 words, ~374 tokens.

Download SKILL.mdSave it as .claude/skills/preprocessing-data-with-automated-pipelines/SKILL.md (or your agent's skills folder). This skill also uses 8 other files; get the full folder from GitHub.
name
preprocessing-data-with-automated-pipelines
description
Design and implement repeatable preprocessing pipelines for cleaning, encoding, transforming, and validating ML input data.
allowed-tools
Read, Write, Edit, Grep, Glob, Bash(cmd:*)
version
1.0.0
author
Jeremy Longshore <jeremy@intentsolutions.io>
license
MIT

Data Preprocessing Pipeline

Positioning

Use this skill as the direct owner for ML input-preparation pipelines.

It covers preprocessing-heavy tasks where the requested deliverable is a repeatable pipeline for cleaning, encoding, transforming, and validating input data.

When to Use

Use this skill when:

  • Prepare raw data for machine learning models.
  • Automate data cleaning and transformation processes.
  • Implement a robust ETL (Extract, Transform, Load) pipeline.

Not For / Boundaries

  • Whole-task ML ownership: use scikit-learn or ml-pipeline-workflow
  • Leakage and prediction-time auditing: use ml-data-leakage-guard
  • Grouped scientific preprocessing with stronger methodological constraints: use scientific-data-preprocessing

Typical Outputs

  • A preprocessing pipeline plan or implementation sketch
  • Clear sequencing for clean, encode, transform, and validate steps
  • Notes that identify where leakage review, training, or evaluation should be run next
  • ml-data-leakage-guard before trusting fitted preprocessing steps
  • splitting-datasets when the next narrow problem is partition strategy

© foryourhealth111-pixel, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 8 other files (scripts, references, assets) in bundled/skills/preprocessing-data-with-automated-pipelines of foryourhealth111-pixel/Vibe-Skills.

  • SKILL.md
  • assets/README.md
  • assets/example_data.csv
  • references/README.md
  • scripts/README.md
  • scripts/handle_errors.py
  • scripts/pipeline.py
  • scripts/transform_data.py
  • scripts/validate_data.py

Open the folder on GitHubat commit ddcaa2a

Compare with similar skills

Preprocessing Data With Automated Pipelines next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Preprocessing Data With Automated Pipelines compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Preprocessing Data With Automated Pipelines this skillforyourhealth111-pixel/Vibe-Skills3.6k—~374Automated safety check: PassMIT
Scikit LearnzLanqing/codex-claude-academic-skills4.7k16 repos~3.9kAutomated safety check: PassBSD-3-Clause
Agentic Kaggle WorkflowFrankS-IntelLab/agentic-kaggle-skill188—~4kAutomated safety check: PassMIT
Senior Data ScientistRaidriar7170/hermes-skilleval1255 repos~1.4kAutomated safety check: PassMIT
Geomlitalo-goncalves/geoML109—~4.9kAutomated safety check: PassGPL-3.0
QuantMind Training Config Generatorqusong0627/QuantMind1.7k—~1.5kAutomated safety check: PassAGPL-3.0

Similar skills

  • Scikit Learn

    zLanqing/codex-claude-academic-skills

    Machine learning in Python with scikit-learn. An agent skill from zLanqing/codex-claude-academic-skills.

    4.7k GitHub starsUsed in 16 repos~3.9k tokens
    Data & AnalyticsAuto-check passed
  • Agentic Kaggle Workflow

    FrankS-IntelLab/agentic-kaggle-skill

    Takes a Kaggle competition from rules and validation design through baselines, ensembling and notebook architecture to a scored submission.

    188 GitHub stars~4k tokensUpdated 3 mo ago
    Data & AnalyticsAuto-check passed
  • Senior Data Scientist

    Raidriar7170/hermes-skilleval

    World-class data science skill for statistical modeling, experimentation, causal inference, and advanced analytics.

    125 GitHub starsUsed in 5 repos~1.4k tokens
    Data & AnalyticsAuto-check passed
  • Geoml

    italo-goncalves/geoML

    Working knowledge of the geoML Python package (github.com/italo-goncalves/geoML): variational Gaussian processes for spatial data, implicit geological modelling, block models, drillhole data…

    109 GitHub stars~4.9k tokensUpdated yesterday
    Data & AnalyticsAuto-check passed
  • Turns a plain-language model training request into a validated QuantMind training config file that can be imported from the Model Training page.

    1.7k GitHub stars~1.5k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Retention Analysis

    liangdabiao/claude-data-analysis-ultra-main

    Analyze user retention and churn using survival analysis, cohort analysis, and machine learning.

    290 GitHub stars~1.3k tokensUpdated 5 mo ago
    Data & AnalyticsAuto-check: notes

More from foryourhealth111-pixel/Vibe-Skills

All 81 skills in this repo
  • Market Research Reports

    foryourhealth111-pixel/Vibe-Skills

    Produces long consulting-style market research and industry reports covering market sizing, competitive landscape, market entry and investment theses.

    3.6k GitHub stars~2.5k tokensUpdated 1 mo ago
    Auto-check: notes
  • Academic Venue Templates

    foryourhealth111-pixel/Vibe-Skills

    Supplies venue-specific LaTeX templates and formatting rules for journals, conferences and posters, and checks a manuscript against page limits and submission requirements.

    3.6k GitHub stars~3.9k tokensUpdated 1 mo ago
    Auto-check: notes
  • Digital Brain

    foryourhealth111-pixel/Vibe-Skills

    This skill should be used when the user asks to "write a post", "check my voice", "look up contact", "prepare for meeting", "weekly review", "track goals", or mentions personal brand, content…

    3.6k GitHub stars~1.7k tokensUpdated 1 mo ago
    Auto-check passed
  • Smart File Writer

    foryourhealth111-pixel/Vibe-Skills

    Diagnoses why a file write failed (permissions, disk space, path length, locks, read-only mounts) before retrying, instead of repeating the same call blindly.

    3.6k GitHub stars~2.6k tokensUpdated 1 mo ago
    Auto-check passed
  • Automated Video Studio

    foryourhealth111-pixel/Vibe-Skills

    Turns footage, audio and a storyboard plan into a finished short video with FFmpeg jump-cuts, subtitle burn-in and a final polish pass.

    3.6k GitHub stars~838 tokensUpdated 1 mo ago
    Auto-check passed
  • Citation Management

    foryourhealth111-pixel/Vibe-Skills

    Turns DOIs, PMIDs and arXiv IDs into clean BibTeX, searches Google Scholar and PubMed, and checks and deduplicates a reference list.

    3.6k GitHub stars~7.6k tokensUpdated 1 mo ago
    Auto-check: notes

Questions about Preprocessing Data With Automated Pipelines

What does Preprocessing Data With Automated Pipelines do?

Design and implement repeatable preprocessing pipelines for cleaning, encoding, transforming, and validating ML input data. Preprocessing Data With Automated Pipelines is an agent skill from foryourhealth111-pixel/Vibe-Skills. Design and implement repeatable preprocessing pipelines for cleaning, encoding, transforming, and validating ML input data.

When should I use Preprocessing Data With Automated Pipelines?

Preprocessing Data With Automated Pipelines fits situations like: tasks that involve Machine learning.

How do I install Preprocessing Data With Automated Pipelines in Claude Code?

Run `npx skills add foryourhealth111-pixel/Vibe-Skills --skill preprocessing-data-with-automated-pipelines -a claude-code`. Or copy the skill folder (bundled/skills/preprocessing-data-with-automated-pipelines in foryourhealth111-pixel/Vibe-Skills) into .claude/skills/preprocessing-data-with-automated-pipelines in your project. Claude Code loads it when a task matches its description.

How do I install Preprocessing Data With Automated Pipelines in Codex?

Run `npx skills add foryourhealth111-pixel/Vibe-Skills --skill preprocessing-data-with-automated-pipelines -a codex`. Or copy the skill folder (bundled/skills/preprocessing-data-with-automated-pipelines in foryourhealth111-pixel/Vibe-Skills) into .agents/skills/preprocessing-data-with-automated-pipelines in your project. Codex loads it when a task matches its description.

Can I use Preprocessing Data With Automated Pipelines in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add foryourhealth111-pixel/Vibe-Skills --skill preprocessing-data-with-automated-pipelines -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/preprocessing-data-with-automated-pipelines, .gemini/skills/preprocessing-data-with-automated-pipelines, .github/skills/preprocessing-data-with-automated-pipelines and .opencode/skills/preprocessing-data-with-automated-pipelines in your project.

What does Preprocessing Data With Automated Pipelines need to run?

Going by SKILL.md and its folder, Preprocessing Data With Automated Pipelines needs Python for the scripts in its folder. Our summary lists: Python 3. Its frontmatter pre-approves these tools: Read, Write, Edit, Grep, Glob, Bash(cmd:*).

Does Preprocessing Data With Automated Pipelines access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Preprocessing Data With Automated Pipelines safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Preprocessing Data With Automated Pipelines use?

Preprocessing Data With Automated Pipelines is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Preprocessing Data With Automated Pipelines use?

About 374 tokens (SKILL.md is roughly 1.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 110 tokens, read only when the agent opens those files.

What are the alternatives to Preprocessing Data With Automated Pipelines?

Skills that share tags, products or a category with Preprocessing Data With Automated Pipelines: Scikit Learn (zLanqing/codex-claude-academic-skills, 4.7k stars), Agentic Kaggle Workflow (FrankS-IntelLab/agentic-kaggle-skill, 188 stars), Senior Data Scientist (Raidriar7170/hermes-skilleval, 125 stars) and Geoml (italo-goncalves/geoML, 109 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Preprocessing Data With Automated Pipelines?

foryourhealth111-pixel (a GitHub user) maintains it in foryourhealth111-pixel/Vibe-Skills, which has 3,627 GitHub stars. The repository holds 81 skills in this directory. The repository was last updated on August 31, 2026.

Source: foryourhealth111-pixel/Vibe-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.