Agent skill

Datarobot Data Preparation

by Kilo-Org in Kilo-Org/kilo-marketplace

Tools and guidance for data upload, dataset management, data validation, and preparing data for DataRobot projects.

Apache-2.0Auto-check passedData & Analytics

Install Datarobot Data Preparation

skills CLI
$ npx skills add Kilo-Org/kilo-marketplace --skill datarobot-data-preparation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Kilo-Org/kilo-marketplace datarobot-data-preparation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Kilo-Org/kilo-marketplace.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/datarobot-data-preparation .claude/skills/datarobot-data-preparation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
datarobot-data-preparation
GitHub stars
190
Token cost
~1.8k tokens
SKILL.md length
655 words
Files
4 (incl. scripts)
Skills in repo
85
Repo updated
First seen
Licence
Apache-2.0

At a glance

Tools and guidance for data upload, dataset management, data validation, and preparing data for DataRobot projects.

  • Works in 4 steps: Dataset Upload → Data Validation → Dataset Management → …
  • Uploading datasets
  • SKILL.md covers Quick Start, When to use this skill, Key capabilities and Workflow examples, plus 9 more sections
  • Runs Python scripts from its folder; calls pip and python; needs DATAROBOT_API_TOKEN

What it does

Datarobot Data Preparation is an agent skill from Kilo-Org/kilo-marketplace. Tools and guidance for data upload, dataset management, data validation, and preparing data for DataRobot projects. Use when uploading datasets, managing data, or validating data for DataRobot.

Its SKILL.md is about 1.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including scripts (for example `scripts/upload_dataset.py`).

It sits in Data & Analytics, covering Data cleaning and CSV and tabular files. The repository describes itself as: Kilo Marketplace - A curated collection of Skills, MCP Servers, and Modes for enhancing AI agent capabilities across the Kilo ecosystem—including Kilo Code (VS Code extension)… The licence is Apache-2.0.

When your agent uses it

  • Uploading datasets
  • Validating data for DataRobot

Example prompts

  • “Use the datarobot-data-preparation skill to tool and guidance for data upload, dataset management, data validation, and preparing data for DataRobot…”
  • “/datarobot-data-preparation”

Requirements

  • Python 3
  • A credential in DATAROBOT_API_TOKEN

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Dataset Upload
  2. Data Validation
  3. Dataset Management
  4. Data Preparation

What it can do on your machine

Read from SKILL.md and the folder at commit ff51758. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • pip
    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • docs.datarobot.com
    • datarobot-public-api-client.readthedocs-hosted.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • DATAROBOT_API_TOKEN

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Datarobot Data Preparation loads about 1.8k tokens when it runs. Until then it costs about 55 tokens; SKILL.md has 655 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~55
When it runs · the whole SKILL.md, loaded when a task matches
~1.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from Kilo-Org/kilo-marketplace at commit ff51758, republished under its Apache-2.0 licence (© Kilo-Org). 655 words, ~1,815 tokens.

Download SKILL.mdSave it as .claude/skills/datarobot-data-preparation/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
datarobot-data-preparation
description
Tools and guidance for data upload, dataset management, data validation, and preparing data for DataRobot projects. Use when uploading datasets, managing data, or validating data for DataRobot.
metadata.category
data

DataRobot Data Preparation Skill

This skill provides guidance for preparing and managing data in DataRobot, including uploading datasets, validating data quality, and managing dataset versions.

Quick Start

Most common use case: Upload and validate a dataset

  1. Upload dataset: upload_dataset(file_path, dataset_name) to upload data
  2. Validate data: validate_dataset(dataset_id) to check data quality
  3. Check schema: get_dataset_schema(dataset_id) to review structure

Example: "Upload sales_data.csv and check if it's ready for training"

When to use this skill

Use this skill when you need to:

  • Upload datasets to DataRobot
  • Validate data before project creation
  • Manage dataset versions and updates
  • Check data quality and completeness
  • Prepare data for training or predictions
  • Handle data format conversions
  • Connect to external data sources

Key capabilities

1. Dataset Upload
  • Upload CSV, Parquet, and other file formats
  • Connect to databases and data warehouses
  • Handle large datasets efficiently
  • Manage dataset metadata and descriptions
2. Data Validation
  • Validate data formats and schemas
  • Check for missing values and data quality issues
  • Verify column types and formats
  • Identify potential data problems
3. Dataset Management
  • List and search datasets
  • Update dataset metadata
  • Create dataset versions
  • Delete or archive old datasets
4. Data Preparation
  • Clean and preprocess data
  • Handle missing values
  • Format data for DataRobot requirements
  • Prepare prediction datasets

Workflow examples

Example 1: Upload and validate dataset

User request: "Upload my sales_data.csv file and check if it's ready for training."

Agent workflow:

  1. Upload the CSV file to DataRobot
  2. Validate the dataset structure and format
  3. Check for missing values and data quality issues
  4. Verify column types are appropriate
  5. Check for potential issues (leakage, formatting)
  6. Report validation results and recommendations
Example 2: Prepare prediction dataset

User request: "Prepare a prediction dataset based on the training data structure from project abc123."

Agent workflow:

  1. Get the training dataset structure from the project
  2. Identify required columns and data types
  3. Create a template with the same structure
  4. Validate the template matches requirements
  5. Provide guidance on filling in prediction values

Using DataRobot SDK

This skill guides you to use the DataRobot Python SDK directly. Install the SDK if needed:

bash
pip install datarobot
Key SDK Operations

Use these DataRobot SDK methods for data management:

Dataset Operations:

  • dr.Dataset.create_from_file(file_path, name) - Upload dataset
  • dr.Dataset.get(dataset_id) - Get dataset details
  • dr.Dataset.list() - List all datasets
  • dataset.row_count - Get row count
  • dataset.column_count - Get column count

Dataset Information:

  • dataset.name - Dataset name
  • dataset.id - Dataset ID
  • dataset.created_at - Creation timestamp

See the Common Patterns section below for complete examples.

Show full SKILL.md (255 more words)Show less

Helper Scripts

This skill includes executable helper scripts that the agent can run directly:

  • scripts/upload_dataset.py - Upload a dataset file to DataRobot

Usage example:

bash
# Upload dataset
python scripts/upload_dataset.py sales_data.csv "Sales Data Q4 2024"

The agent can run this script directly or use it as reference when writing code.

Best practices

  1. Data quality: Clean and validate data before upload
  2. File formats: Use appropriate formats (CSV for small, Parquet for large)
  3. Naming conventions: Use clear, descriptive dataset names
  4. Metadata: Add descriptions and tags for better organization
  5. Versioning: Create versions for important datasets
  6. Data validation: Always validate data before using in projects

Common patterns

Pattern 1: Upload and validate
python
import datarobot as dr
import os

# Initialize client
client = dr.Client(
    token=os.getenv("DATAROBOT_API_TOKEN"),
    endpoint=os.getenv("DATAROBOT_ENDPOINT")
)

# Upload dataset
dataset = dr.Dataset.create_from_file(
    file_path="sales_data.csv",
    name="Sales Data Q4 2024"
)

print(f"Dataset ID: {dataset.id}")
print(f"Rows: {dataset.row_count}, Columns: {dataset.column_count}")

# Get dataset details
dataset_info = dr.Dataset.get(dataset.id)
print(f"Dataset name: {dataset_info.name}")
print(f"Created: {dataset_info.created_at}")
Pattern 2: Dataset management
python
import datarobot as dr

# List all datasets
datasets = dr.Dataset.list()
print(f"Found {len(datasets)} datasets")

# Search for specific dataset
for dataset in datasets:
    if "sales" in dataset.name.lower():
        print(f"Found: {dataset.name} (ID: {dataset.id})")

# Get specific dataset
dataset = dr.Dataset.get("abc123")
print(f"Dataset: {dataset.name}")
print(f"Size: {dataset.row_count} rows x {dataset.column_count} columns")

Data format requirements

CSV Files
  • UTF-8 encoding recommended
  • Headers in first row
  • Consistent delimiters (comma, tab)
  • Proper date/time formatting
Parquet Files
  • Columnar format, efficient for large datasets
  • Preserves data types
  • Better compression than CSV
Database Connections
  • Support for various databases
  • Connection credentials required
  • Query-based data access

Data quality checks

Common checks to perform:

  • Missing values: Identify columns with high missing value rates
  • Data types: Verify columns have correct types
  • Value ranges: Check for outliers and invalid values
  • Duplicates: Identify duplicate records
  • Consistency: Check for data consistency issues

Error handling

Common errors and solutions:

  • Upload failures: Check file format, size limits, encoding
  • Validation errors: Fix data quality issues before proceeding
  • Schema mismatches: Ensure data structure matches expectations
  • Access issues: Verify permissions for dataset operations

SDK Setup

Install DataRobot SDK
bash
pip install datarobot
Initialize Client
python
import datarobot as dr
import os

client = dr.Client(
    token=os.getenv("DATAROBOT_API_TOKEN"),
    endpoint=os.getenv("DATAROBOT_ENDPOINT", "https://app.datarobot.com")
)

Resources

© Kilo-Org, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (scripts) in skills/datarobot-data-preparation of Kilo-Org/kilo-marketplace.

  • SKILL.md
  • LICENSE
  • local.patch
  • scripts/upload_dataset.py

Open the folder on GitHubat commit ff51758

Compare with similar skills

Datarobot Data Preparation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Datarobot Data Preparation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Datarobot Data Preparation this skillKilo-Org/kilo-marketplace190—~1.8kAutomated safety check: PassApache-2.0
Verified Data Analysis with pandaspipeshub-ai/pipeshub-ai3.8k—~1.2kAutomated safety check: PassApache-2.0
Portaljs Check Data Qualitydatopian/portaljs2.4k1 repos~1.5kAutomated safety check: PassMIT
Dataset Quality Auditzebbern/claude-code-guide4.7k—~996Automated safety check: PassMIT
Splitting Datasetsjeremylongshore/tons-of-skills-marketplace2.8k1 repos~836Automated safety check: PassMIT
Data Table AnalysisNVIDIA-AI-Blueprints/deep-researcher-agent883—~2.5kAutomated safety check: PassApache-2.0

Similar skills

  • Verified Data Analysis with pandas

    pipeshub-ai/pipeshub-ai

    Loads, cleans, aggregates and joins tabular data with pandas under a verification rule: every number reported must be one that the code actually printed.

    3.8k GitHub stars~1.2k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Audit a local or remote tabular file (CSV/TSV) for common data quality issues — schema, nulls, types, duplicates.

    2.4k GitHub starsUsed in 1 repo~1.5k tokens
    Data & AnalyticsAuto-check passed
  • Dataset Quality Audit

    zebbern/claude-code-guide

    Run comprehensive quality checks on tabular data (CSV/Excel/TSV/JSON), detecting missing values, duplicates, outliers, format issues, and type inconsistencies to produce an overall score, grade, and…

    4.7k GitHub stars~996 tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Splitting Datasets

    jeremylongshore/tons-of-skills-marketplace

    Process split datasets into training, validation, and testing sets for ML model development.

    2.8k GitHub starsUsed in 1 repo~836 tokens
    Data & AnalyticsAuto-check passed
  • Data Table Analysis

    NVIDIA-AI-Blueprints/deep-researcher-agent

    A skill your agent uses for converting researched facts or user-provided data into structured tables by writing code, then running Python/pandas calculations in the job-scoped sandbox.

    883 GitHub stars~2.5k tokensUpdated 2 days ago
    Data & AnalyticsAuto-check passed
  • A skill your agent uses to migrate course catalog data from external sources (CSV, PDF, website) and bulk-create Learning and LearningCourse records in Education Cloud.

    1.1k GitHub stars~5.4k tokensUpdated yesterday
    Data & AnalyticsAuto-check passed

More from Kilo-Org/kilo-marketplace

All 85 skills in this repo
  • AzureML Project Scaffolding

    Kilo-Org/kilo-marketplace

    Sets up and maintains AzureML-ready Python projects as uv workspaces with devcontainers, a Makefile and job YAML, so local runs match cloud jobs and experiments stay reproducible.

    190 GitHub stars~3.1k tokensUpdated 10 days ago
    Auto-check: notes
  • Jupyter Notebook Builder

    Kilo-Org/kilo-marketplace

    Creates, inspects, edits and runs Jupyter notebooks, scaffolding experiment or tutorial notebooks from templates and preferring a Jupyter MCP server over raw JSON edits.

    190 GitHub stars~1.3k tokensUpdated 10 days ago
    Auto-check passed
  • Tableau Dashboard Creator

    Kilo-Org/kilo-marketplace

    Takes a plain-language dashboard request through brand setup, data exploration, planning, an interactive HTML mock and a Tableau implementation spec.

    190 GitHub stars~3.8k tokensUpdated 10 days ago
    Auto-check: notes
  • Elasticsearch File Ingest

    Kilo-Org/kilo-marketplace

    Ingest and transform data files (CSV/JSON/Parquet/Arrow IPC) into Elasticsearch with stream processing and custom transforms.

    190 GitHub stars~2.8k tokensUpdated 10 days ago
    Auto-check passed
  • Nifi Flow Layout

    Kilo-Org/kilo-marketplace

    A skill your agent uses when arranging Apache NiFi processors, process groups, ports, comments, numbering, crossing connections, dense fan-in/fan-out, or reusable readable canvas layouts.

    190 GitHub stars~1.5k tokensUpdated 10 days ago
    Auto-check passed
  • Splunk Ingest Processor Setup

    Kilo-Org/kilo-marketplace

    Render Cisco Data Fabric ingest-time routing workflows and Splunk Cloud Platform Ingest Processor setup plans with SPL2 pipelines, source types, destinations, lifecycle handoffs, queue and…

    190 GitHub stars~1.2k tokensUpdated 10 days ago
    Auto-check passed

Questions about Datarobot Data Preparation

What does Datarobot Data Preparation do?

Tools and guidance for data upload, dataset management, data validation, and preparing data for DataRobot projects. Datarobot Data Preparation is an agent skill from Kilo-Org/kilo-marketplace. Tools and guidance for data upload, dataset management, data validation, and preparing data for DataRobot projects.

When should I use Datarobot Data Preparation?

Datarobot Data Preparation fits situations like: uploading datasets; validating data for DataRobot.

How do I install Datarobot Data Preparation in Claude Code?

Run `npx skills add Kilo-Org/kilo-marketplace --skill datarobot-data-preparation -a claude-code`. Or copy the skill folder (skills/datarobot-data-preparation in Kilo-Org/kilo-marketplace) into .claude/skills/datarobot-data-preparation in your project. Claude Code loads it when a task matches its description.

How do I install Datarobot Data Preparation in Codex?

Run `npx skills add Kilo-Org/kilo-marketplace --skill datarobot-data-preparation -a codex`. Or copy the skill folder (skills/datarobot-data-preparation in Kilo-Org/kilo-marketplace) into .agents/skills/datarobot-data-preparation in your project. Codex loads it when a task matches its description.

Can I use Datarobot Data Preparation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Kilo-Org/kilo-marketplace --skill datarobot-data-preparation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/datarobot-data-preparation, .gemini/skills/datarobot-data-preparation, .github/skills/datarobot-data-preparation and .opencode/skills/datarobot-data-preparation in your project.

What does Datarobot Data Preparation need to run?

Going by SKILL.md and its folder, Datarobot Data Preparation needs Python for the scripts in its folder, the command-line tools its instructions call (pip and python) and credentials named DATAROBOT_API_TOKEN. Our summary lists: Python 3; A credential in DATAROBOT_API_TOKEN.

Does Datarobot Data Preparation access the network?

SKILL.md names 2 domains. As links in the text: docs.datarobot.com and datarobot-public-api-client.readthedocs-hosted.com. This is read from the text; nothing was executed.

Is Datarobot Data Preparation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Datarobot Data Preparation use?

Datarobot Data Preparation is published under the Apache-2.0 licence (from the LICENSE file in the skill folder). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Datarobot Data Preparation use?

About 1.8k tokens (SKILL.md is roughly 7.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Datarobot Data Preparation?

Skills that share tags, products or a category with Datarobot Data Preparation: Verified Data Analysis with pandas (pipeshub-ai/pipeshub-ai, 3.8k stars), Portaljs Check Data Quality (datopian/portaljs, 2.4k stars), Dataset Quality Audit (zebbern/claude-code-guide, 4.7k stars) and Splitting Datasets (jeremylongshore/tons-of-skills-marketplace, 2.8k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Datarobot Data Preparation?

Kilo-Org (a GitHub organization) maintains it in Kilo-Org/kilo-marketplace, which has 190 GitHub stars. The repository holds 85 skills in this directory. The repository was last updated on September 28, 2026.

Source: Kilo-Org/kilo-marketplace on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.