Split datasets into training, validation, and test partitions with the right stratification and temporal rules.

MITAuto-check passedData & Analytics

Install Splitting Datasets

skills CLI
$ npx skills add foryourhealth111-pixel/Vibe-Skills --skill splitting-datasets -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install foryourhealth111-pixel/Vibe-Skills splitting-datasets --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/foryourhealth111-pixel/Vibe-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/bundled/skills/splitting-datasets .claude/skills/splitting-datasets && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
splitting-datasets
GitHub stars
3.6k
Token cost
~346 tokens
SKILL.md length
110 words
Files
8 (incl. scripts, references, assets)
Skills in repo
81
Repo updated
First seen
Licence
MIT

At a glance

Split datasets into training, validation, and test partitions with the right stratification and temporal rules.

  • Tasks that involve Data cleaning
  • SKILL.md covers Positioning, When to Use, Not For / Boundaries and Typical Outputs, plus 1 more section
  • Runs Python scripts from its folder

What it does

Splitting Datasets is an agent skill from foryourhealth111-pixel/Vibe-Skills. Split datasets into training, validation, and test partitions with the right stratification and temporal rules. Use as a narrow preprocessing helper once the broader ML workflow is already chosen, not as the main route owner for an end-to-end ML task.

Its SKILL.md is about 350 tokens, which your agent loads only when the skill is triggered. The skill folder holds 10 other files, including scripts, reference files and assets (for example `assets/README.md`, `assets/dataset_schema.json` and `assets/split_data_config.yaml`).

It sits in Data & Analytics, covering Data cleaning. The repository describes itself as: Intelligent Skill routing and workflow orchestration for AI agents — +21.12 pp reward, −29.6% tokens on SkillsBench with DeepSeekV4Flash-VE. The licence is MIT.

When your agent uses it

  • Tasks that involve Data cleaning

Example prompts

  • “/splitting-datasets”

Requirements

  • Python 3
  • Pre-approved tools (allowed-tools): Read, Write, Edit, Grep, Glob, Bash(cmd:*)

What it can do on your machine

Read from SKILL.md and the folder at commit ddcaa2a. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Write
    • Edit
    • Grep
    • Glob
    • Bash(cmd:*)

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Splitting Datasets loads about 346 tokens when it runs, and up to ~478 if it reads all its reference files. Until then it costs about 68 tokens; SKILL.md has 110 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~68
When it runs · the whole SKILL.md, loaded when a task matches
~346
With references · SKILL.md plus every file in references/, read only if the agent opens them
~478

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from foryourhealth111-pixel/Vibe-Skills at commit ddcaa2a, republished under its MIT licence (© foryourhealth111-pixel). 110 words, ~346 tokens.

Download SKILL.mdSave it as .claude/skills/splitting-datasets/SKILL.md (or your agent's skills folder). This skill also uses 7 other files; get the full folder from GitHub.
name
splitting-datasets
description
Split datasets into training, validation, and test partitions with the right stratification and temporal rules. Use as a narrow preprocessing helper once the broader ML workflow is already chosen, not as the main route owner for an end-to-end ML task.
allowed-tools
Read, Write, Edit, Grep, Glob, Bash(cmd:*)
version
1.0.0
author
Jeremy Longshore <jeremy@intentsolutions.io>
license
MIT

Dataset Splitter

Positioning

Treat this skill as a narrow helper for partition strategy.

When to Use

Use this skill when:

  • Prepare a dataset for machine learning model training.
  • Create training, validation, and testing sets.
  • Partition data to evaluate model performance.

Not For / Boundaries

  • Full preprocessing-pipeline ownership: use preprocessing-data-with-automated-pipelines
  • Leakage audits and prediction-time checks: use ml-data-leakage-guard
  • Model training and tuning after the split: use scikit-learn

Typical Outputs

  • Partition strategy with ratios, random seeds, and stratification rules
  • Notes on temporal or grouped split constraints
  • Handoff guidance for leakage review and downstream training
  • preprocessing-data-with-automated-pipelines for the broader preprocessing sequence
  • ml-data-leakage-guard to verify the split does not leak future or test information

© foryourhealth111-pixel, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 7 other files (scripts, references, assets) in bundled/skills/splitting-datasets of foryourhealth111-pixel/Vibe-Skills.

  • SKILL.md
  • assets/README.md
  • assets/dataset_schema.json
  • assets/example_dataset.csv
  • assets/split_data_config.yaml
  • references/README.md
  • scripts/README.md
  • scripts/split_data.py

Open the folder on GitHubat commit ddcaa2a

Compare with similar skills

Splitting Datasets next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Splitting Datasets compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Splitting Datasets this skillforyourhealth111-pixel/Vibe-Skills3.6k—~346Automated safety check: PassMIT
Question2reportrefraction-ray/xalpha2.7k—~3.2kAutomated safety check: PassMIT
Dingo VerifyMigoXLab/dingo757—~741Automated safety check: NotesApache-2.0
Data Validationplatonai/Browser41.2k—~896Automated safety check: PassApache-2.0
Pandas ProJeffallan/claude-skills12k1 repos~1.5kAutomated safety check: PassMIT
Issues DeduplicationJetBrains/ideavim10k—~1.3kAutomated safety check: PassMIT

Similar skills

  • Question2report

    refraction-ray/xalpha

    Turn a natural-language financial question into a polished, self-contained HTML report.

    2.7k GitHub stars~3.2k tokensUpdated 2 mo ago
    Data & AnalyticsAuto-check passed
  • Dingo Verify

    MigoXLab/dingo

    A skill your agent uses when the user wants to fact-check an article or verify factual claims in a document.

    757 GitHub stars~741 tokensUpdated yesterday
    Data & AnalyticsAuto-check: notes
  • Data Validation

    platonai/Browser4

    Validates data against common and custom rules (required fields, formats, ranges).

    1.2k GitHub stars~896 tokensUpdated yesterday
    Data & AnalyticsAuto-check passed
  • Pandas Pro

    Jeffallan/claude-skills

    Handles pandas DataFrame work: cleaning, merging, groupby aggregation, pivots, time-series resampling and memory tuning, with checks on dtypes, shapes and nulls.

    12k GitHub starsUsed in 1 repo~1.5k tokens
    Data & AnalyticsAuto-check passed
  • Issues Deduplication

    JetBrains/ideavim

    Official

    Handles deduplication of YouTrack issues. An agent skill from JetBrains/ideavim.

    10k GitHub stars~1.3k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Openbb Data Fetcher

    monarchjuno/vibe-investing

    Fetch financial, market, economic, fundamental, news, options, crypto, ETF, index, and macro data through the OpenBB Python interface instead of the OpenBB MCP server.

    299 GitHub stars~2.9k tokensUpdated 5 mo ago
    Data & AnalyticsAuto-check: notes

More from foryourhealth111-pixel/Vibe-Skills

All 81 skills in this repo
  • Market Research Reports

    foryourhealth111-pixel/Vibe-Skills

    Produces long consulting-style market research and industry reports covering market sizing, competitive landscape, market entry and investment theses.

    3.6k GitHub stars~2.5k tokensUpdated 1 mo ago
    Auto-check: notes
  • Academic Venue Templates

    foryourhealth111-pixel/Vibe-Skills

    Supplies venue-specific LaTeX templates and formatting rules for journals, conferences and posters, and checks a manuscript against page limits and submission requirements.

    3.6k GitHub stars~3.9k tokensUpdated 1 mo ago
    Auto-check: notes
  • Digital Brain

    foryourhealth111-pixel/Vibe-Skills

    This skill should be used when the user asks to "write a post", "check my voice", "look up contact", "prepare for meeting", "weekly review", "track goals", or mentions personal brand, content…

    3.6k GitHub stars~1.7k tokensUpdated 1 mo ago
    Auto-check passed
  • Smart File Writer

    foryourhealth111-pixel/Vibe-Skills

    Diagnoses why a file write failed (permissions, disk space, path length, locks, read-only mounts) before retrying, instead of repeating the same call blindly.

    3.6k GitHub stars~2.6k tokensUpdated 1 mo ago
    Auto-check passed
  • Automated Video Studio

    foryourhealth111-pixel/Vibe-Skills

    Turns footage, audio and a storyboard plan into a finished short video with FFmpeg jump-cuts, subtitle burn-in and a final polish pass.

    3.6k GitHub stars~838 tokensUpdated 1 mo ago
    Auto-check passed
  • Citation Management

    foryourhealth111-pixel/Vibe-Skills

    Turns DOIs, PMIDs and arXiv IDs into clean BibTeX, searches Google Scholar and PubMed, and checks and deduplicates a reference list.

    3.6k GitHub stars~7.6k tokensUpdated 1 mo ago
    Auto-check: notes

Questions about Splitting Datasets

What does Splitting Datasets do?

Split datasets into training, validation, and test partitions with the right stratification and temporal rules. Splitting Datasets is an agent skill from foryourhealth111-pixel/Vibe-Skills. Split datasets into training, validation, and test partitions with the right stratification and temporal rules.

When should I use Splitting Datasets?

Splitting Datasets fits situations like: tasks that involve Data cleaning.

How do I install Splitting Datasets in Claude Code?

Run `npx skills add foryourhealth111-pixel/Vibe-Skills --skill splitting-datasets -a claude-code`. Or copy the skill folder (bundled/skills/splitting-datasets in foryourhealth111-pixel/Vibe-Skills) into .claude/skills/splitting-datasets in your project. Claude Code loads it when a task matches its description.

How do I install Splitting Datasets in Codex?

Run `npx skills add foryourhealth111-pixel/Vibe-Skills --skill splitting-datasets -a codex`. Or copy the skill folder (bundled/skills/splitting-datasets in foryourhealth111-pixel/Vibe-Skills) into .agents/skills/splitting-datasets in your project. Codex loads it when a task matches its description.

Can I use Splitting Datasets in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add foryourhealth111-pixel/Vibe-Skills --skill splitting-datasets -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/splitting-datasets, .gemini/skills/splitting-datasets, .github/skills/splitting-datasets and .opencode/skills/splitting-datasets in your project.

What does Splitting Datasets need to run?

Going by SKILL.md and its folder, Splitting Datasets needs Python for the scripts in its folder. Our summary lists: Python 3. Its frontmatter pre-approves these tools: Read, Write, Edit, Grep, Glob, Bash(cmd:*).

Does Splitting Datasets access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Splitting Datasets safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Splitting Datasets use?

Splitting Datasets is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Splitting Datasets use?

About 346 tokens (SKILL.md is roughly 1.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 132 tokens, read only when the agent opens those files.

What are the alternatives to Splitting Datasets?

Skills that share tags, products or a category with Splitting Datasets: Question2report (refraction-ray/xalpha, 2.7k stars), Dingo Verify (MigoXLab/dingo, 757 stars), Data Validation (platonai/Browser4, 1.2k stars) and Pandas Pro (Jeffallan/claude-skills, 12k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Splitting Datasets?

foryourhealth111-pixel (a GitHub user) maintains it in foryourhealth111-pixel/Vibe-Skills, which has 3,627 GitHub stars. The repository holds 81 skills in this directory. The repository was last updated on August 31, 2026.

Source: foryourhealth111-pixel/Vibe-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.