Agent skill

Pandas Data Wrangling

by wentorai in wentorai/research-plugins

Data cleaning, transformation, and exploratory analysis with pandas

MITAuto-check passedData & Analytics

Install Pandas Data Wrangling

skills CLI
$ npx skills add wentorai/research-plugins --skill pandas-data-wrangling -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install wentorai/research-plugins pandas-data-wrangling --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/wentorai/research-plugins.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/analysis/wrangling/pandas-data-wrangling .claude/skills/pandas-data-wrangling && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
pandas-data-wrangling
GitHub stars
298
Used in
1 other repo
Token cost
~2k tokens
SKILL.md length
346 words
Files
1
Skills in repo
405
Repo updated
First seen
Licence
MIT

At a glance

Data cleaning, transformation, and exploratory analysis with pandas

  • Tasks that involve Data cleaning
  • SKILL.md covers Overview, Loading and Inspecting Data, Handling Missing Data and Data Transformation, plus 4 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Tasks that involve DataFrames

What it does

Pandas Data Wrangling is an agent skill from wentorai/research-plugins. Data cleaning, transformation, and exploratory analysis with pandas

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Data & Analytics, covering Data cleaning and DataFrames. It works with pandas. The repository describes itself as: 350+ academic research skills, MCP configs, and plugins for Research-Claw and AI agents. The licence is MIT.

When your agent uses it

  • Tasks that involve Data cleaning
  • Tasks that involve DataFrames

Example prompts

  • “/pandas-data-wrangling”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit bf44b3c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • pandas.pydata.org
    • wesmckinney.com
    • store.metasnake.com
    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Pandas Data Wrangling loads about 2k tokens when it runs. Until then it costs about 22 tokens; SKILL.md has 346 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~22
When it runs · the whole SKILL.md, loaded when a task matches
~2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from wentorai/research-plugins at commit bf44b3c, republished under its MIT licence (© wentorai). 346 words, ~1,952 tokens.

Download SKILL.mdSave it as .claude/skills/pandas-data-wrangling/SKILL.md (or your agent's skills folder).
name
pandas-data-wrangling
description
Data cleaning, transformation, and exploratory analysis with pandas

Pandas Data Wrangling Guide

Overview

Data wrangling -- the process of cleaning, transforming, and preparing raw data for analysis -- typically consumes 60-80% of a data scientist's time. Pandas is the de facto standard library for tabular data manipulation in Python, and mastering its idioms directly translates to faster, more reliable research workflows.

This guide covers the essential pandas operations that researchers encounter daily: loading heterogeneous data sources, diagnosing data quality issues, handling missing values, reshaping data for analysis, and performing exploratory data analysis (EDA). Each section includes copy-paste code examples designed for real-world research datasets.

Whether you are cleaning survey responses, preprocessing experimental logs, merging datasets from multiple sources, or preparing features for machine learning, the patterns here will save hours of trial and error.

Loading and Inspecting Data

Reading Common Formats
python
import pandas as pd
import numpy as np

# CSV with encoding and date parsing
df = pd.read_csv('data.csv', encoding='utf-8',
                 parse_dates=['timestamp'],
                 dtype={'participant_id': str})

# Excel with specific sheet
df = pd.read_excel('data.xlsx', sheet_name='Experiment1',
                   header=1)  # Skip first row

# JSON (nested)
df = pd.json_normalize(json_data, record_path='results',
                       meta=['experiment_id', 'date'])

# Parquet (fast, columnar)
df = pd.read_parquet('data.parquet')
Initial Diagnostics
python
# Shape and types
print(f"Shape: {df.shape}")
print(df.dtypes)
print(df.info(memory_usage='deep'))

# Statistical summary
print(df.describe(include='all'))

# Missing value report
missing = df.isnull().sum()
missing_pct = (missing / len(df) * 100).round(1)
missing_report = pd.DataFrame({
    'count': missing,
    'percent': missing_pct
}).query('count > 0').sort_values('percent', ascending=False)
print(missing_report)

# Duplicate check
n_dupes = df.duplicated().sum()
print(f"Duplicate rows: {n_dupes}")

Handling Missing Data

Strategy Decision Tree
SituationStrategypandas Method
< 5% missing, randomDrop rowsdf.dropna()
Numeric, moderate missingMean/median imputationdf.fillna(df.median())
Categorical missingMode or "Unknown"df.fillna('Unknown')
Time series gapsForward/backward filldf.ffill() / df.bfill()
Systematic missingMultiple imputationsklearn.impute.IterativeImputer
Feature > 50% missingDrop columndf.drop(columns=[...])
Implementation Examples
python
# Conditional imputation
df['age'] = df['age'].fillna(df.groupby('group')['age'].transform('median'))

# Interpolation for time series
df['temperature'] = df['temperature'].interpolate(method='time')

# Flag missing values before imputing (preserve information)
df['salary_missing'] = df['salary'].isnull().astype(int)
df['salary'] = df['salary'].fillna(df['salary'].median())

Data Transformation

Type Conversion and Cleaning
python
# String cleaning
df['name'] = df['name'].str.strip().str.lower()
df['email'] = df['email'].str.replace(r'\s+', '', regex=True)

# Categorical conversion (saves memory, enables ordering)
df['education'] = pd.Categorical(
    df['education'],
    categories=['high_school', 'bachelors', 'masters', 'phd'],
    ordered=True
)

# Numeric extraction from text
df['value'] = df['text_field'].str.extract(r'(\d+\.?\d*)').astype(float)
Reshaping Operations
python
# Wide to long (unpivot)
df_long = pd.melt(df,
    id_vars=['subject_id', 'condition'],
    value_vars=['score_t1', 'score_t2', 'score_t3'],
    var_name='timepoint',
    value_name='score'
)

# Long to wide (pivot)
df_wide = df_long.pivot_table(
    index='subject_id',
    columns='condition',
    values='score',
    aggfunc='mean'
).reset_index()

# Cross-tabulation
ct = pd.crosstab(df['group'], df['outcome'],
                 margins=True, normalize='index')
Merging and Joining
python
# Left join with validation
merged = pd.merge(
    experiments, participants,
    on='participant_id',
    how='left',
    validate='many_to_one',  # Catch unexpected duplicates
    indicator=True           # Shows _merge column
)

# Check merge quality
print(merged['_merge'].value_counts())

Exploratory Data Analysis (EDA)

Automated EDA Pipeline
python
def quick_eda(df, target_col=None):
    """Run a quick EDA pipeline on a DataFrame."""
    print(f"=== Shape: {df.shape} ===\n")

    # Numeric columns
    numeric_cols = df.select_dtypes(include=np.number).columns
    print(f"Numeric columns ({len(numeric_cols)}):")
    print(df[numeric_cols].describe().round(2))

    # Categorical columns
    cat_cols = df.select_dtypes(include=['object', 'category']).columns
    print(f"\nCategorical columns ({len(cat_cols)}):")
    for col in cat_cols:
        n_unique = df[col].nunique()
        print(f"  {col}: {n_unique} unique values")
        if n_unique <= 10:
            print(f"    {df[col].value_counts().to_dict()}")

    # Correlations with target
    if target_col and target_col in numeric_cols:
        corr = df[numeric_cols].corr()[target_col].drop(target_col)
        print(f"\nCorrelations with '{target_col}':")
        print(corr.sort_values(ascending=False).round(3))

quick_eda(df, target_col='accuracy')
GroupBy Aggregations
python
# Multi-metric summary by group
summary = df.groupby('method').agg(
    mean_acc=('accuracy', 'mean'),
    std_acc=('accuracy', 'std'),
    median_time=('runtime_sec', 'median'),
    n_runs=('run_id', 'count')
).round(3).sort_values('mean_acc', ascending=False)

print(summary.to_markdown())

Performance Optimization

TechniqueWhen to UseSpeedup
pd.Categorical for stringsRepeated string values2-10x memory
.query() instead of boolean indexingComplex filters1.5-3x
pd.eval() for arithmeticColumn arithmetic2-5x
Parquet instead of CSVLarge datasets5-20x I/O
df.pipe() for chainingReadable pipelinesClarity
python
# Method chaining with pipe
result = (
    df
    .query('score > 0')
    .assign(log_score=lambda x: np.log1p(x['score']))
    .groupby('group')
    .agg(mean_log=('log_score', 'mean'))
    .sort_values('mean_log', ascending=False)
)

Best Practices

  • Never modify the original DataFrame in place. Use .copy() when creating derived datasets.
  • Use method chaining for readability. Pipe operations together instead of creating intermediate variables.
  • Document your cleaning steps. Keep a data cleaning log or use a Jupyter notebook with explanations.
  • Validate after every merge. Check row counts, null values, and the _merge indicator column.
  • Profile before optimizing. Use df.memory_usage(deep=True) to identify memory bottlenecks.
  • Save intermediate results as Parquet. It preserves dtypes and is much faster than CSV.

References

© wentorai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/analysis/wrangling/pandas-data-wrangling of wentorai/research-plugins.

Open the folder on GitHubat commit bf44b3c

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in wentorai/research-plugins, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Pandas Data Wrangling next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Pandas Data Wrangling compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Pandas Data Wrangling this skillwentorai/research-plugins2981 repos~2kAutomated safety check: PassMIT
Pandas ProJeffallan/claude-skills12k1 repos~1.5kAutomated safety check: PassMIT
Verified Data Analysis with pandaspipeshub-ai/pipeshub-ai3.8k—~1.2kAutomated safety check: PassApache-2.0
Data Table AnalysisNVIDIA-AI-Blueprints/deep-researcher-agent886—~2.5kAutomated safety check: PassApache-2.0
Data TransformMicrock/ordinary-claude-skills4043 repos~4.3kAutomated safety check: PassCustom licence
CSV Processingbenchflow-ai/skillsbench1.8k—~455Automated safety check: PassApache-2.0

Similar skills

  • Pandas Pro

    Jeffallan/claude-skills

    Handles pandas DataFrame work: cleaning, merging, groupby aggregation, pivots, time-series resampling and memory tuning, with checks on dtypes, shapes and nulls.

    12k GitHub starsUsed in 1 repo~1.5k tokens
    Data & AnalyticsAuto-check passed
  • Verified Data Analysis with pandas

    pipeshub-ai/pipeshub-ai

    Loads, cleans, aggregates and joins tabular data with pandas under a verification rule: every number reported must be one that the code actually printed.

    3.8k GitHub stars~1.2k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Data Table Analysis

    NVIDIA-AI-Blueprints/deep-researcher-agent

    A skill your agent uses for converting researched facts or user-provided data into structured tables by writing code, then running Python/pandas calculations in the job-scoped sandbox.

    886 GitHub stars~2.5k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Data Transform

    Microck/ordinary-claude-skills

    Transform, clean, reshape, and preprocess data using pandas and numpy.

    404 GitHub starsUsed in 3 repos~4.3k tokens
    Data & AnalyticsAuto-check passed
  • CSV Processing

    benchflow-ai/skillsbench

    A skill your agent uses when reading sensor data from CSV files, writing simulation results to CSV, processing time-series data with pandas, or handling missing values in datasets.

    1.8k GitHub stars~455 tokensUpdated 2 mo ago
    Data & AnalyticsAuto-check passed
  • Data Cleaning

    ericrisco/rsc-harness

    A skill your agent uses when a raw table is too dirty to trust — nulls, sentinels, duplicate rows, category sprawl, mixed types, bad dates — and you need a re-runnable clean() plus a schema gate…

    180 GitHub stars~3.6k tokensUpdated today
    Data & AnalyticsAuto-check passed

More from wentorai/research-plugins

All 405 skills in this repo
  • Abstract Writing Guide

    wentorai/research-plugins

    Craft structured research abstracts that maximize clarity and journal acceptance

    298 GitHub starsUsed in 1 repo~1.7k tokens
    Auto-check passed
  • Academic Citation Manager

    wentorai/research-plugins

    Manage academic citations across BibTeX, APA, MLA, and Chicago formats

    298 GitHub starsUsed in 1 repo~2.7k tokens
    Auto-check passed
  • Academic Paper Summarizer

    wentorai/research-plugins

    Summarize academic papers with structured extraction of key elements

    298 GitHub starsUsed in 1 repo~1.4k tokens
    Auto-check passed
  • Academic Study Methods

    wentorai/research-plugins

    Evidence-based study techniques for academic learning and retention

    298 GitHub starsUsed in 1 repo~1.8k tokens
    Auto-check passed
  • Academic Tone Guide

    wentorai/research-plugins

    Adjust writing tone and register for academic audiences and venues

    298 GitHub starsUsed in 1 repo~1.9k tokens
    Auto-check passed
  • Academic Translation Guide

    wentorai/research-plugins

    Academic translation, post-editing, and Chinglish correction guide

    298 GitHub starsUsed in 1 repo~1.6k tokens
    Auto-check passed

Works with

Questions about Pandas Data Wrangling

What does Pandas Data Wrangling do?

Data cleaning, transformation, and exploratory analysis with pandas. Pandas Data Wrangling is an agent skill from wentorai/research-plugins.

When should I use Pandas Data Wrangling?

Pandas Data Wrangling fits situations like: tasks that involve Data cleaning; tasks that involve DataFrames.

How do I install Pandas Data Wrangling in Claude Code?

Run `npx skills add wentorai/research-plugins --skill pandas-data-wrangling -a claude-code`. Or copy the skill folder (skills/analysis/wrangling/pandas-data-wrangling in wentorai/research-plugins) into .claude/skills/pandas-data-wrangling in your project. Claude Code loads it when a task matches its description.

How do I install Pandas Data Wrangling in Codex?

Run `npx skills add wentorai/research-plugins --skill pandas-data-wrangling -a codex`. Or copy the skill folder (skills/analysis/wrangling/pandas-data-wrangling in wentorai/research-plugins) into .agents/skills/pandas-data-wrangling in your project. Codex loads it when a task matches its description.

Can I use Pandas Data Wrangling in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add wentorai/research-plugins --skill pandas-data-wrangling -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/pandas-data-wrangling, .gemini/skills/pandas-data-wrangling, .github/skills/pandas-data-wrangling and .opencode/skills/pandas-data-wrangling in your project.

What does Pandas Data Wrangling need to run?

SKILL.md names no scripts, command-line tools or credentials: Pandas Data Wrangling is instructions for the agent only. Our summary lists: Python 3.

Does Pandas Data Wrangling access the network?

SKILL.md names 4 domains. As links in the text: pandas.pydata.org, wesmckinney.com, store.metasnake.com and github.com. This is read from the text; nothing was executed.

Is Pandas Data Wrangling safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Pandas Data Wrangling use?

Pandas Data Wrangling is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Pandas Data Wrangling use?

About 2k tokens (SKILL.md is roughly 7.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Pandas Data Wrangling?

Skills that share tags, products or a category with Pandas Data Wrangling: Pandas Pro (Jeffallan/claude-skills, 12k stars), Verified Data Analysis with pandas (pipeshub-ai/pipeshub-ai, 3.8k stars), Data Table Analysis (NVIDIA-AI-Blueprints/deep-researcher-agent, 886 stars) and Data Transform (Microck/ordinary-claude-skills, 404 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Pandas Data Wrangling?

wentorai (a GitHub user) maintains it in wentorai/research-plugins, which has 298 GitHub stars. The repository holds 405 skills in this directory. The repository was last updated on June 19, 2026.

Source: wentorai/research-plugins on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.