Process split datasets into training, validation, and testing sets for ML model development.

MITAuto-check passedData & Analytics

Install Splitting Datasets

skills CLI
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill splitting-datasets -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jeremylongshore/tons-of-skills-marketplace splitting-datasets --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/.curated/splitting-datasets .claude/skills/splitting-datasets && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
splitting-datasets
GitHub stars
2.8k
Token cost
~836 tokens
SKILL.md length
394 words
Files
7 (incl. scripts, references, assets)
Skills in repo
3,342
Repo updated
First seen
Licence
MIT

At a glance

Process split datasets into training, validation, and testing sets for ML model development.

  • Works in 3 steps: Analyze Request: The skill analyzes the… → Generate Code: Based on the request, the… → Execute Splitting: The code is executed…
  • Requesting split dataset
  • SKILL.md covers Overview, How It Works, When to Use This Skill and Examples, plus 7 more sections
  • Train-test split

What it does

Splitting Datasets is an agent skill from jeremylongshore/tons-of-skills-marketplace. Process split datasets into training, validation, and testing sets for ML model development. Use when requesting "split dataset", "train-test split", or "data partitioning". Trigger with relevant phrases based on skill purpose.

Its SKILL.md is about 840 tokens, which your agent loads only when the skill is triggered. The skill folder holds 9 other files, including scripts, reference files and assets (for example `assets/README.md`, `assets/dataset_schema.json` and `assets/split_data_config.yaml`). Compatibility notes: Designed for Claude Code

It sits in Data & Analytics, covering Data cleaning, Machine learning and CSV and tabular files. It works with Python. The repository describes itself as: Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com. The licence is MIT.

When your agent uses it

  • Requesting split dataset
  • Train-test split
  • Data partitioning
  • With relevant phrases based on skill purpose

Example prompts

  • “split dataset”
  • “train-test split”
  • “data partitioning”
  • “/splitting-datasets”

Requirements

  • Python 3
  • Compatibility (from SKILL.md): Designed for Claude Code
  • Pre-approved tools (allowed-tools): Read, Write, Edit, Grep, Glob, Bash(cmd:*)

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Analyze Request: The skill analyzes the user's request to determine the dataset to be split and the desired proportions for each subset.
  2. Generate Code: Based on the request, the skill generates Python code utilizing standard ML libraries to perform the data splitting.
  3. Execute Splitting: The code is executed to split the dataset into training, validation, and testing sets according to the specified ratios.

What it can do on your machine

Read from SKILL.md and the folder at commit cfae287. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Write
    • Edit
    • Grep
    • Glob
    • Bash(cmd:*)

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/, which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Designed for Claude Code

    From compatibility in the SKILL.md frontmatter.

Context cost

Splitting Datasets loads about 836 tokens when it runs, and up to ~851 if it reads all its reference files. Until then it costs about 62 tokens; SKILL.md has 394 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~62
When it runs · the whole SKILL.md, loaded when a task matches
~836
With references · SKILL.md plus every file in references/, read only if the agent opens them
~851

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from jeremylongshore/tons-of-skills-marketplace at commit cfae287, republished under its MIT licence (© jeremylongshore). 394 words, ~836 tokens.

Download SKILL.mdSave it as .claude/skills/splitting-datasets/SKILL.md (or your agent's skills folder). This skill also uses 6 other files; get the full folder from GitHub.
name
splitting-datasets
description
Process split datasets into training, validation, and testing sets for ML model development. Use when requesting "split dataset", "train-test split", or "data partitioning". Trigger with relevant phrases based on skill purpose.
allowed-tools
Read, Write, Edit, Grep, Glob, Bash(cmd:*)
compatibility
Designed for Claude Code
version
1.22.0
author
Jeremy Longshore <jeremy@intentsolutions.io>
license
MIT
tags
ai, testing, ml

Dataset Splitter

Split datasets into training, validation, and testing sets with configurable ratios and stratification options.

Overview

This skill automates the process of dividing a dataset into subsets for training, validating, and testing machine learning models. It ensures proper data preparation and facilitates robust model evaluation.

How It Works

  1. Analyze Request: The skill analyzes the user's request to determine the dataset to be split and the desired proportions for each subset.
  2. Generate Code: Based on the request, the skill generates Python code utilizing standard ML libraries to perform the data splitting.
  3. Execute Splitting: The code is executed to split the dataset into training, validation, and testing sets according to the specified ratios.

When to Use This Skill

This skill activates when you need to:

  • Prepare a dataset for machine learning model training.
  • Create training, validation, and testing sets.
  • Partition data to evaluate model performance.

Examples

Example 1: Splitting a CSV file

User request: "Split the data in 'my_data.csv' into 70% training, 15% validation, and 15% testing sets."

The skill will:

  1. Generate Python code to read the 'my_data.csv' file.
  2. Execute the code to split the data according to the specified proportions, creating 'train.csv', 'validation.csv', and 'test.csv' files.
Example 2: Creating a Train-Test Split

User request: "Create a train-test split of 'large_dataset.csv' with an 80/20 ratio."

The skill will:

  1. Generate Python code to load 'large_dataset.csv'.
  2. Execute the code to split the dataset into 80% training and 20% testing sets, saving them as 'train.csv' and 'test.csv'.
Show full SKILL.md (144 more words)Show less

Best Practices

  • Data Integrity: Verify that the splitting process maintains the integrity of the data, ensuring no data loss or corruption.
  • Stratification: Consider stratification when splitting imbalanced datasets to maintain class distributions in each subset.
  • Randomization: Ensure the splitting process is randomized to avoid bias in the resulting datasets.

Integration

This skill can be integrated with other data processing and model training tools within the Claude Code ecosystem to create a complete machine learning workflow.

Prerequisites

  • Appropriate file access permissions
  • Required dependencies installed

Instructions

  1. Invoke this skill when the trigger conditions are met
  2. Provide necessary context and parameters
  3. Review the generated output
  4. Apply modifications as needed

Output

The skill produces structured output relevant to the task.

Error Handling

  • Invalid input: Prompts for correction
  • Missing dependencies: Lists required components
  • Permission errors: Suggests remediation steps

Resources

  • Project documentation
  • Related skills and commands

© jeremylongshore, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 6 other files (scripts, references, assets) in skills/.curated/splitting-datasets of jeremylongshore/tons-of-skills-marketplace.

  • SKILL.md
  • assets/README.md
  • assets/dataset_schema.json
  • assets/example_dataset.csv
  • assets/split_data_config.yaml
  • references/README.md
  • scripts/README.md

Open the folder on GitHubat commit cfae287

Compare with similar skills

Splitting Datasets next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Splitting Datasets compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Splitting Datasets this skilljeremylongshore/tons-of-skills-marketplace2.8k—~836Automated safety check: PassMIT
Data Table AnalysisNVIDIA-AI-Blueprints/deep-researcher-agent886—~2.5kAutomated safety check: PassApache-2.0
Sap Hana Cloud Data Intelligencesecondsky/sap-skills462—~3.2kAutomated safety check: PassGPL-3.0
Empirical Analysis Skill PythonDrchronx/ai-agent-research-starter-kit139—~3kAutomated safety check: PassCustom licence
XLSX Spreadsheet ToolkitXiaomiMiMo/MiMo-Code14k—~2.9kAutomated safety check: PassApache-2.0
CSV and Excel MergerOneWave-AI/claude-skills336—~1.6kAutomated safety check: PassMIT

Similar skills

  • Data Table Analysis

    NVIDIA-AI-Blueprints/deep-researcher-agent

    A skill your agent uses for converting researched facts or user-provided data into structured tables by writing code, then running Python/pandas calculations in the job-scoped sandbox.

    886 GitHub stars~2.5k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Develops data processing pipelines, integrations, and machine learning scenarios in SAP Data Intelligence Cloud.

    462 GitHub stars~3.2k tokensUpdated 5 days ago
    Data & AnalyticsAuto-check passed
  • Empirical Analysis Skill Python

    Drchronx/ai-agent-research-starter-kit

    Parameterized Python empirical-analysis and machine-learning workflow for applied economics, public health epidemiology, supervised ML, and ML causal inference.

    139 GitHub stars~3k tokensUpdated 4 mo ago
    Data & AnalyticsAuto-check passed
  • XLSX Spreadsheet Toolkit

    XiaomiMiMo/MiMo-Code

    Builds, edits, cleans, recalculates and reads Excel workbooks and CSV files with openpyxl and pandas, plus LibreOffice for recalculation and PDF export.

    14k GitHub stars~2.9k tokensUpdated yesterday
    Documents & OfficeAuto-check passed
  • CSV and Excel Merger

    OneWave-AI/claude-skills

    Combines CSV, TSV and Excel files into one verified table with pandas, by stacking or joining, mapping columns, normalizing keys and removing duplicates.

    336 GitHub stars~1.6k tokensUpdated 8 days ago
    Documents & OfficeAuto-check passed
  • Official

    Creates, edits and analyzes spreadsheets (.xlsx, .xlsm, .csv, .tsv) with openpyxl and pandas, writing live formulas and recalculating to confirm zero formula errors.

    180k GitHub starsUsed in 4 repos~2.1k tokens
    Documents & OfficeAuto-check passed

More from jeremylongshore/tons-of-skills-marketplace

All 3,342 skills in this repo
  • Performing Security Code Review

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to conduct a security-focused code review using the security-agent plugin.

    2.8k GitHub starsUsed in 2 repos~1.3k tokens
    Auto-check: notes
  • Adapting Transfer Learning Models

    jeremylongshore/tons-of-skills-marketplace

    Build this skill automates the adaptation of pre-trained machine learning models using transfer learning techniques.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Agent Context Loader

    jeremylongshore/tons-of-skills-marketplace

    Execute proactive auto-loading: automatically detects and loads agents.md files.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Aggregating Performance Metrics

    jeremylongshore/tons-of-skills-marketplace

    Aggregate and centralize performance metrics from applications, systems, databases, caches, and services.

    2.8k GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Analyzing Capacity Planning

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to analyze capacity requirements and plan for future growth.

    2.8k GitHub stars~947 tokensUpdated today
    Auto-check passed
  • Analyzing Database Indexes

    jeremylongshore/tons-of-skills-marketplace

    Process use when you need to work with database indexing. An agent skill from jeremylongshore/tons-of-skills-marketplace.

    2.8k GitHub stars~2k tokensUpdated today
    Auto-check passed

Works with

Questions about Splitting Datasets

What does Splitting Datasets do?

Process split datasets into training, validation, and testing sets for ML model development. Splitting Datasets is an agent skill from jeremylongshore/tons-of-skills-marketplace. Process split datasets into training, validation, and testing sets for ML model development.

When should I use Splitting Datasets?

Splitting Datasets fits situations like: requesting split dataset; train-test split; data partitioning; with relevant phrases based on skill purpose.

How do I install Splitting Datasets in Claude Code?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill splitting-datasets -a claude-code`. Or copy the skill folder (skills/.curated/splitting-datasets in jeremylongshore/tons-of-skills-marketplace) into .claude/skills/splitting-datasets in your project. Claude Code loads it when a task matches its description.

How do I install Splitting Datasets in Codex?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill splitting-datasets -a codex`. Or copy the skill folder (skills/.curated/splitting-datasets in jeremylongshore/tons-of-skills-marketplace) into .agents/skills/splitting-datasets in your project. Codex loads it when a task matches its description.

Can I use Splitting Datasets in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill splitting-datasets -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/splitting-datasets, .gemini/skills/splitting-datasets, .github/skills/splitting-datasets and .opencode/skills/splitting-datasets in your project.

What does Splitting Datasets need to run?

SKILL.md names no scripts, command-line tools or credentials: Splitting Datasets is instructions for the agent only. Our summary lists: Python 3. Its frontmatter pre-approves these tools: Read, Write, Edit, Grep, Glob, Bash(cmd:*). Compatibility (from SKILL.md): Designed for Claude Code.

Does Splitting Datasets access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Splitting Datasets safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Splitting Datasets use?

Splitting Datasets is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Splitting Datasets use?

About 836 tokens (SKILL.md is roughly 3.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 15 tokens, read only when the agent opens those files.

What are the alternatives to Splitting Datasets?

Skills that share tags, products or a category with Splitting Datasets: Data Table Analysis (NVIDIA-AI-Blueprints/deep-researcher-agent, 886 stars), Sap Hana Cloud Data Intelligence (secondsky/sap-skills, 462 stars), Empirical Analysis Skill Python (Drchronx/ai-agent-research-starter-kit, 139 stars) and XLSX Spreadsheet Toolkit (XiaomiMiMo/MiMo-Code, 14k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Splitting Datasets?

jeremylongshore (a GitHub user) maintains it in jeremylongshore/tons-of-skills-marketplace, which has 2,827 GitHub stars. The repository holds 3,342 skills in this directory. The repository was last updated on October 10, 2026.

Source: jeremylongshore/tons-of-skills-marketplace on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.