Agent skill

Pandas Pro

by Jeffallan in Jeffallan/claude-skills

Handles pandas DataFrame work: cleaning, merging, groupby aggregation, pivots, time-series resampling and memory tuning, with checks on dtypes, shapes and nulls.

MITAuto-check passedData & Analytics

Install Pandas Pro

skills CLI
$ npx skills add Jeffallan/claude-skills --skill pandas-pro -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Jeffallan/claude-skills pandas-pro --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Jeffallan/claude-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/pandas-pro .claude/skills/pandas-pro && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
pandas-pro
GitHub stars
12k
Used in
1 other repo
Token cost
~1.5k tokens
SKILL.md length
295 words
Files
6 (incl. references)
Skills in repo
58
Repo updated
First seen
Licence
MIT

At a glance

Handles pandas DataFrame work: cleaning, merging, groupby aggregation, pivots, time-series resampling and memory tuning, with checks on dtypes, shapes and nulls.

  • Works in 5 steps: Assess data structure — Examine dtypes,… → Design transformation — Plan vectorized… → Implement efficiently — Use vectorized… → …
  • Cleaning a messy DataFrame with missing values, duplicates and wrong types
  • SKILL.md covers Core Workflow, Reference Guide, Code Patterns and Constraints, plus 1 more section
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

The agent inspects dtypes, memory usage and missing values first, plans vectorized operations in place of loops, then implements with method chaining and proper indexing. Results are checked with assertions on row counts, null counts and expected columns. For large frames it profiles memory, applies categorical types and reads in chunks when needed.

Code patterns show swapping iterrows loops for vectorized math, avoiding chained indexing by taking a copy, groupby with named aggregations, merges on several keys with validation, forward-fill followed by interpolation for gaps, and resampling of time series. Reference files cover DataFrame operations, data cleaning, aggregation and groupby, merging and joining, and performance.

When your agent uses it

  • Cleaning a messy DataFrame with missing values, duplicates and wrong types
  • Joining DataFrames on several keys and checking the result
  • Building pivot tables and groupby summaries
  • Resampling and filling gaps in time-series data
  • Cutting memory use on a large dataset

Example prompts

  • “Merge orders.csv and customers.csv on customer_id and date, then check that no rows were duplicated.”
  • “Replace this iterrows loop with vectorized pandas code.”
  • “Resample the sensor readings to daily means and interpolate the missing days.”
  • “This DataFrame uses too much memory, so convert suitable columns to categoricals and read the file in chunks.”

Requirements

  • Python with pandas

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Assess data structure — Examine dtypes, memory usage, missing values, data quality
  2. Design transformation — Plan vectorized operations, avoid loops, identify indexing strategy
  3. Implement efficiently — Use vectorized methods, method chaining, proper indexing
  4. Validate results — Check dtypes, shapes, null counts, and row counts
  5. Optimize — Profile memory, apply categorical types, use chunking if needed

What it can do on your machine

Read from SKILL.md and the folder at commit 1be15d8. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com
    • synergetic.solutions
    • jeffallan.github.io

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Pandas Pro loads about 1.5k tokens when it runs, and up to ~18k if it reads all its reference files. Until then it costs about 117 tokens; SKILL.md has 295 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~117
When it runs · the whole SKILL.md, loaded when a task matches
~1.5k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~18k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Jeffallan/claude-skills at commit 1be15d8, republished under its MIT licence (© Jeffallan). 295 words, ~1,539 tokens.

Download SKILL.mdSave it as .claude/skills/pandas-pro/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
pandas-pro
description
Performs pandas DataFrame operations for data analysis, manipulation, and transformation. Use when working with pandas DataFrames, data cleaning, aggregation, merging, or time series analysis. Invoke for data manipulation tasks such as joining DataFrames on multiple keys, pivoting tables, resampling time series, handling NaN values with interpolation or forward-fill, groupby aggregations, type conversion, or performance optimization of large datasets.
license
MIT
metadata.author
https://github.com/Jeffallan
metadata.company
https://synergetic.solutions
metadata.version
1.1.0
metadata.domain
data-ml
metadata.triggers
pandas, DataFrame, data manipulation, data cleaning, aggregation, groupby, merge, join, time series, data wrangling, pivot table, data transformation
metadata.role
expert
metadata.scope
implementation
metadata.output-format
code
metadata.related-skills
python-pro

Pandas Pro

Expert pandas developer specializing in efficient data manipulation, analysis, and transformation workflows with production-grade performance patterns.

Core Workflow

  1. Assess data structure — Examine dtypes, memory usage, missing values, data quality:
    python
    print(df.dtypes)
    print(df.memory_usage(deep=True).sum() / 1e6, "MB")
    print(df.isna().sum())
    print(df.describe(include="all"))
  2. Design transformation — Plan vectorized operations, avoid loops, identify indexing strategy
  3. Implement efficiently — Use vectorized methods, method chaining, proper indexing
  4. Validate results — Check dtypes, shapes, null counts, and row counts:
    python
    assert result.shape[0] == expected_rows, f"Row count mismatch: {result.shape[0]}"
    assert result.isna().sum().sum() == 0, "Unexpected nulls after transform"
    assert set(result.columns) == expected_cols
  5. Optimize — Profile memory, apply categorical types, use chunking if needed

Reference Guide

Load detailed guidance based on context:

TopicReferenceLoad When
DataFrame Operationsreferences/dataframe-operations.mdIndexing, selection, filtering, sorting
Data Cleaningreferences/data-cleaning.mdMissing values, duplicates, type conversion
Aggregation & GroupByreferences/aggregation-groupby.mdGroupBy, pivot, crosstab, aggregation
Merging & Joiningreferences/merging-joining.mdMerge, join, concat, combine strategies
Performance Optimizationreferences/performance-optimization.mdMemory usage, vectorization, chunking

Code Patterns

Vectorized Operations (before/after)
python
# ❌ AVOID: row-by-row iteration
for i, row in df.iterrows():
    df.at[i, 'tax'] = row['price'] * 0.2

# ✅ USE: vectorized assignment
df['tax'] = df['price'] * 0.2
Safe Subsetting with .copy()
python
# ❌ AVOID: chained indexing triggers SettingWithCopyWarning
df['A']['B'] = 1

# ✅ USE: .loc[] with explicit copy when mutating a subset
subset = df.loc[df['status'] == 'active', :].copy()
subset['score'] = subset['score'].fillna(0)
GroupBy Aggregation
python
summary = (
    df.groupby(['region', 'category'], observed=True)
    .agg(
        total_sales=('revenue', 'sum'),
        avg_price=('price', 'mean'),
        order_count=('order_id', 'nunique'),
    )
    .reset_index()
)
Merge with Validation
python
merged = pd.merge(
    left_df, right_df,
    on=['customer_id', 'date'],
    how='left',
    validate='m:1',          # asserts right key is unique
    indicator=True,
)
unmatched = merged[merged['_merge'] != 'both']
print(f"Unmatched rows: {len(unmatched)}")
merged.drop(columns=['_merge'], inplace=True)
Missing Value Handling
python
# Forward-fill then interpolate numeric gaps
df['price'] = df['price'].ffill().interpolate(method='linear')

# Fill categoricals with mode, numerics with median
for col in df.select_dtypes(include='object'):
    df[col] = df[col].fillna(df[col].mode()[0])
for col in df.select_dtypes(include='number'):
    df[col] = df[col].fillna(df[col].median())
Time Series Resampling
python
daily = (
    df.set_index('timestamp')
    .resample('D')
    .agg({'revenue': 'sum', 'sessions': 'count'})
    .fillna(0)
)
Pivot Table
python
pivot = df.pivot_table(
    values='revenue',
    index='region',
    columns='product_line',
    aggfunc='sum',
    fill_value=0,
    margins=True,
)
Memory Optimization
python
# Downcast numerics and convert low-cardinality strings to categorical
df['category'] = df['category'].astype('category')
df['count'] = pd.to_numeric(df['count'], downcast='integer')
df['score'] = pd.to_numeric(df['score'], downcast='float')
print(df.memory_usage(deep=True).sum() / 1e6, "MB after optimization")

Constraints

MUST DO
  • Use vectorized operations instead of loops
  • Set appropriate dtypes (categorical for low-cardinality strings)
  • Check memory usage with .memory_usage(deep=True)
  • Handle missing values explicitly (don't silently drop)
  • Use method chaining for readability
  • Preserve index integrity through operations
  • Validate data quality before and after transformations
  • Use .copy() when modifying subsets to avoid SettingWithCopyWarning
MUST NOT DO
  • Iterate over DataFrame rows with .iterrows() unless absolutely necessary
  • Use chained indexing (df['A']['B']) — use .loc[] or .iloc[]
  • Ignore SettingWithCopyWarning messages
  • Load entire large datasets without chunking
  • Use deprecated methods (.ix, .append() — use pd.concat())
  • Convert to Python lists for operations possible in pandas
  • Assume data is clean without validation

Output Templates

When implementing pandas solutions, provide:

  1. Code with vectorized operations and proper indexing
  2. Comments explaining complex transformations
  3. Memory/performance considerations if dataset is large
  4. Data validation checks (dtypes, nulls, shapes)

Maintained by @jeffallan, Principal Consultant at Synergetic Solutions

Documentation

© Jeffallan, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files (references) in skills/pandas-pro of Jeffallan/claude-skills.

  • SKILL.md
  • references/aggregation-groupby.md
  • references/data-cleaning.md
  • references/dataframe-operations.md
  • references/merging-joining.md
  • references/performance-optimization.md

Open the folder on GitHubat commit 1be15d8

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in Jeffallan/claude-skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Pandas Pro next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Pandas Pro compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Pandas Pro this skillJeffallan/claude-skills12k1 repos~1.5kAutomated safety check: PassMIT
Data Table AnalysisNVIDIA-AI-Blueprints/deep-researcher-agent886—~2.5kAutomated safety check: PassApache-2.0
Python Executorcortega26/chile-hub1132 repos~1.5kAutomated safety check: PassMIT
Verified Data Analysis with pandaspipeshub-ai/pipeshub-ai3.8k—~1.2kAutomated safety check: PassApache-2.0
Analytics Data AnalysisMindrally/skills271—~1.6kAutomated safety check: PassApache-2.0
CSV and Excel MergerOneWave-AI/claude-skills336—~1.6kAutomated safety check: PassMIT

Similar skills

  • Data Table Analysis

    NVIDIA-AI-Blueprints/deep-researcher-agent

    A skill your agent uses for converting researched facts or user-provided data into structured tables by writing code, then running Python/pandas calculations in the job-scoped sandbox.

    886 GitHub stars~2.5k tokensUpdated yesterday
    Data & AnalyticsAuto-check passed
  • Python Executor

    cortega26/chile-hub

    Execute Python code in a safe sandboxed environment via [inference.sh](https://inference.sh).

    113 GitHub starsUsed in 2 repos~1.5k tokens
    Data & AnalyticsAuto-check passed
  • Verified Data Analysis with pandas

    pipeshub-ai/pipeshub-ai

    Loads, cleans, aggregates and joins tabular data with pandas under a verification rule: every number reported must be one that the code actually printed.

    3.8k GitHub stars~1.2k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Analytics Data Analysis

    Mindrally/skills

    Best practices for analytics, data analysis, and visualization using Python, pandas, matplotlib, seaborn, and Jupyter notebooks.

    271 GitHub stars~1.6k tokensUpdated yesterday
    Data & AnalyticsAuto-check passed
  • CSV and Excel Merger

    OneWave-AI/claude-skills

    Combines CSV, TSV and Excel files into one verified table with pandas, by stacking or joining, mapping columns, normalizing keys and removing duplicates.

    336 GitHub stars~1.6k tokensUpdated 8 days ago
    Documents & OfficeAuto-check passed
  • Chdb Datastore

    vemetric/vemetric

    A skill your agent uses when the user has tabular data (pandas DataFrame, parquet, csv, Arrow, json) and wants to filter, group, aggregate, join, or speed up slow pandas.

    395 GitHub starsUsed in 2 repos~1.4k tokens
    Data & AnalyticsAuto-check passed

More from Jeffallan/claude-skills

All 58 skills in this repo
  • API Designer

    Jeffallan/claude-skills

    Designs REST and GraphQL APIs from resource modeling to an OpenAPI 3.1 contract, with versioning, pagination and RFC 7807 error handling.

    12k GitHub starsUsed in 1 repo~2k tokens
    Auto-check passed
  • CLI Developer

    Jeffallan/claude-skills

    Walks through designing, building and polishing a command-line tool: user workflow and command hierarchy, implementation in commander, click, typer or cobra, completions and cross-platform testing.

    12k GitHub starsUsed in 1 repo~1.2k tokens
    Auto-check passed
  • Kubernetes Specialist

    Jeffallan/claude-skills

    Creates and checks Kubernetes manifests, Helm charts, RBAC and network policies, and helps debug pod problems, with kubectl checks and rollback steps.

    12k GitHub starsUsed in 1 repo~2.1k tokens
    Auto-check passed
  • Laravel Specialist

    Jeffallan/claude-skills

    Builds Laravel 10+ applications with Eloquent models, Sanctum authentication, Horizon queues, API resources and Livewire components, tested with Pest or PHPUnit.

    12k GitHub starsUsed in 1 repo~2.1k tokens
    Auto-check passed
  • Apache Spark Engineer

    Jeffallan/claude-skills

    Guides writing and tuning Apache Spark jobs: DataFrame and RDD code, Spark SQL, partitioning, caching, shuffle tuning and structured streaming.

    12k GitHub starsUsed in 1 repo~1.7k tokens
    Auto-check passed
  • TypeScript Pro

    Jeffallan/claude-skills

    Designs advanced TypeScript types: generics, conditional and mapped types, branded types, discriminated unions and type guards, with tsc checks and tRPC type safety.

    12k GitHub starsUsed in 1 repo~1.4k tokens
    Auto-check passed

Works with

Questions about Pandas Pro

What does Pandas Pro do?

Handles pandas DataFrame work: cleaning, merging, groupby aggregation, pivots, time-series resampling and memory tuning, with checks on dtypes, shapes and nulls. The agent inspects dtypes, memory usage and missing values first, plans vectorized operations in place of loops, then implements with method chaining and proper indexing. Results are checked with assertions on row counts, null counts and expected columns.

When should I use Pandas Pro?

Pandas Pro fits situations like: cleaning a messy DataFrame with missing values, duplicates and wrong types; joining DataFrames on several keys and checking the result; building pivot tables and groupby summaries; resampling and filling gaps in time-series data.

How do I install Pandas Pro in Claude Code?

Run `npx skills add Jeffallan/claude-skills --skill pandas-pro -a claude-code`. Or copy the skill folder (skills/pandas-pro in Jeffallan/claude-skills) into .claude/skills/pandas-pro in your project. Claude Code loads it when a task matches its description.

How do I install Pandas Pro in Codex?

Run `npx skills add Jeffallan/claude-skills --skill pandas-pro -a codex`. Or copy the skill folder (skills/pandas-pro in Jeffallan/claude-skills) into .agents/skills/pandas-pro in your project. Codex loads it when a task matches its description.

Can I use Pandas Pro in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Jeffallan/claude-skills --skill pandas-pro -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/pandas-pro, .gemini/skills/pandas-pro, .github/skills/pandas-pro and .opencode/skills/pandas-pro in your project.

What does Pandas Pro need to run?

SKILL.md names no scripts, command-line tools or credentials: Pandas Pro is instructions for the agent only. Our summary lists: Python with pandas.

Does Pandas Pro access the network?

SKILL.md names 3 domains. As links in the text: github.com, synergetic.solutions and jeffallan.github.io. This is read from the text; nothing was executed.

Is Pandas Pro safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Pandas Pro use?

Pandas Pro is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Pandas Pro use?

About 1.5k tokens (SKILL.md is roughly 6.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 16k tokens, read only when the agent opens those files.

What are the alternatives to Pandas Pro?

Skills that share tags, products or a category with Pandas Pro: Data Table Analysis (NVIDIA-AI-Blueprints/deep-researcher-agent, 886 stars), Python Executor (cortega26/chile-hub, 113 stars), Verified Data Analysis with pandas (pipeshub-ai/pipeshub-ai, 3.8k stars) and Analytics Data Analysis (Mindrally/skills, 271 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Pandas Pro?

Jeffallan (a GitHub user) maintains it in Jeffallan/claude-skills, which has 11,802 GitHub stars. The repository holds 58 skills in this directory. The repository was last updated on October 3, 2026.

Source: Jeffallan/claude-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.