Agent skill

Polars Dataframes

by jaechang-hits in jaechang-hits/SciAgent-Skills

Fast in-memory DataFrame with lazy evaluation, parallel execution, Arrow backend.

MITAuto-check passedData & Analytics

Install Polars Dataframes

skills CLI
$ npx skills add jaechang-hits/SciAgent-Skills --skill polars-dataframes -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jaechang-hits/SciAgent-Skills polars-dataframes --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/scientific-computing/polars-dataframes .claude/skills/polars-dataframes && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
polars-dataframes
GitHub stars
370
Token cost
~6.4k tokens
SKILL.md length
1,391 words
Files
4 (incl. references)
Skills in repo
163
Repo updated
First seen
Licence
MIT

At a glance

Fast in-memory DataFrame with lazy evaluation, parallel execution, Arrow backend.

  • Works in 10 steps: DataFrame Operations → GroupBy & Aggregations → Joins → …
  • Tabular data in RAM (1–100 GB) when pandas is too slow
  • SKILL.md covers Overview, When to Use, Prerequisites and Quick Start, plus 8 more sections
  • Calls pip

What it does

Polars Dataframes is an agent skill from jaechang-hits/SciAgent-Skills. Fast in-memory DataFrame with lazy evaluation, parallel execution, Arrow backend. Use for tabular data in RAM (1–100 GB) when pandas is too slow. Expression API: select, filter, groupby, joins, pivots, window. Lazy mode enables predicate/projection pushdown. Reads CSV, Parquet, JSON, Excel, DBs, cloud. Larger-than-RAM: Dask; GPU: cuDF.

Its SKILL.md is about 6.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including reference files (for example `references/advanced_operations.md`, `references/io_best_practices.md` and `references/pandas_migration.md`).

It sits in Data & Analytics, covering DataFrames. It works with Polars, Dask, pandas and Microsoft Excel. The repository describes itself as: 197 bioinformatics & life science skills for Claude Code and AI agents — BixBench 92.0% accuracy. RNA-seq, single-cell, drug discovery, proteomics, and more. Powers OmicsHorizon. The licence is MIT.

When your agent uses it

  • Tabular data in RAM (1–100 GB) when pandas is too slow
  • Tasks that involve DataFrames

Example prompts

  • “/polars-dataframes”

Requirements

  • Python 3

Workflow steps

10 steps, taken from the step headings in SKILL.md.

  1. DataFrame Operations
  2. GroupBy & Aggregations
  3. Joins
  4. Reshaping
  5. Data I/O
  6. Expression API
  7. Lazy Evaluation
  8. ETL Pipeline (CSV → Clean → Parquet)
  9. Multi-Source Join and Aggregation
  10. Time-Series Feature Engineering

What it can do on your machine

Read from SKILL.md and the folder at commit 82c862c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • docs.pola.rs
    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Polars Dataframes loads about 6.4k tokens when it runs, and up to ~18k if it reads all its reference files. Until then it costs about 89 tokens; SKILL.md has 1,391 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~89
When it runs · the whole SKILL.md, loaded when a task matches
~6.4k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~18k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jaechang-hits/SciAgent-Skills at commit 82c862c, republished under its MIT licence (© jaechang-hits). 1,391 words, ~6,401 tokens.

Download SKILL.mdSave it as .claude/skills/polars-dataframes/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
polars-dataframes
description
Fast in-memory DataFrame with lazy evaluation, parallel execution, Arrow backend. Use for tabular data in RAM (1–100 GB) when pandas is too slow. Expression API: select, filter, group_by, joins, pivots, window. Lazy mode enables predicate/projection pushdown. Reads CSV, Parquet, JSON, Excel, DBs, cloud. Larger-than-RAM: Dask; GPU: cuDF.
license
MIT

Polars DataFrames

Overview

Polars is a high-performance DataFrame library for Python built on Apache Arrow with a Rust backend. It provides an expression-based API with lazy evaluation and automatic parallelization for efficient data processing, transformation, and analysis.

When to Use

  • Processing tabular datasets from 100 MB to 100 GB that fit in RAM
  • ETL pipelines requiring fast read/transform/write cycles
  • Replacing pandas when performance matters (10–100x speedup typical)
  • Lazy query pipelines with automatic optimization (predicate/projection pushdown)
  • Joining, pivoting, and reshaping large tables
  • Reading Parquet, CSV, JSON, or cloud-stored data efficiently
  • Window functions and complex grouped aggregations
  • For larger-than-RAM data, use Dask or Vaex instead
  • For GPU-accelerated DataFrames, use cuDF instead

Prerequisites

bash
pip install polars
# Optional extras:
pip install polars[all]          # All I/O backends
pip install polars[pandas]       # Pandas interop
pip install polars[numpy]        # NumPy interop
pip install connectorx sqlalchemy  # Database connectivity

Quick Start

python
import polars as pl

# Create DataFrame
df = pl.DataFrame({
    "name": ["Alice", "Bob", "Charlie", "Diana"],
    "dept": ["Sales", "Eng", "Sales", "Eng"],
    "salary": [70000, 85000, 72000, 90000],
})

# Expression-based pipeline
result = (
    df.filter(pl.col("salary") > 71000)
    .with_columns(bonus=pl.col("salary") * 0.1)
    .group_by("dept")
    .agg(
        pl.col("salary").mean().alias("avg_salary"),
        pl.len().alias("count"),
    )
)
print(result)
# shape: (2, 3)
# ┌───────┬────────────┬───────┐
# │ dept  ┆ avg_salary ┆ count │
# ├───────┼────────────┼───────┤
# │ Eng   ┆ 87500.0    ┆ 2     │
# │ Sales ┆ 72000.0    ┆ 1     │
# └───────┴────────────┴───────┘

Core API

1. DataFrame Operations

Select, filter, add/modify columns, sort, and sample rows.

python
import polars as pl

df = pl.DataFrame({
    "id": [1, 2, 3, 4, 5],
    "name": ["Alice", "Bob", "Charlie", "Diana", "Eve"],
    "age": [25, 30, 35, 28, 32],
    "score": [88.5, 92.0, 76.3, 95.1, 84.7],
})

# Select columns (with computed expressions)
selected = df.select(
    "name",
    pl.col("age"),
    (pl.col("score") / 100).alias("score_pct"),
)
print(selected.shape)  # (5, 3)

# Filter rows (multiple conditions → implicit AND)
filtered = df.filter(
    pl.col("age") > 27,
    pl.col("score") > 80,
)
print(filtered.shape)  # (3, 4) — Bob, Diana, Eve

# Add columns (preserves existing)
enriched = df.with_columns(
    grade=pl.when(pl.col("score") >= 90).then(pl.lit("A"))
           .when(pl.col("score") >= 80).then(pl.lit("B"))
           .otherwise(pl.lit("C")),
    age_months=pl.col("age") * 12,
)
print(enriched.columns)
# ['id', 'name', 'age', 'score', 'grade', 'age_months']

# Sort
df.sort("score", descending=True).head(3)
2. GroupBy & Aggregations

Group rows and compute summary statistics.

python
import polars as pl

sales = pl.DataFrame({
    "region": ["East", "West", "East", "West", "East", "West"],
    "product": ["A", "A", "B", "B", "A", "B"],
    "revenue": [100, 150, 200, 180, 120, 210],
    "units": [10, 15, 20, 18, 12, 21],
})

# Basic group_by
summary = sales.group_by("region").agg(
    pl.col("revenue").sum().alias("total_rev"),
    pl.col("revenue").mean().alias("avg_rev"),
    pl.len().alias("n_transactions"),
)
print(summary)

# Multiple keys + conditional aggregation
by_rp = sales.group_by("region", "product").agg(
    pl.col("revenue").sum(),
    (pl.col("units") > 15).sum().alias("large_orders"),
)
print(by_rp)
python
# Window functions with over() — add group stats without collapsing rows
enriched = sales.with_columns(
    region_avg=pl.col("revenue").mean().over("region"),
    rank_in_region=pl.col("revenue").rank(descending=True).over("region"),
    pct_of_region=pl.col("revenue") / pl.col("revenue").sum().over("region"),
)
print(enriched.select("region", "product", "revenue", "region_avg", "rank_in_region"))
3. Joins

Combine DataFrames on shared keys.

python
import polars as pl

customers = pl.DataFrame({
    "cid": [1, 2, 3, 4],
    "name": ["Alice", "Bob", "Charlie", "Diana"],
})
orders = pl.DataFrame({
    "oid": [101, 102, 103, 104],
    "cid": [1, 2, 1, 5],
    "amount": [100, 200, 150, 300],
})

# Inner join — only matching rows
inner = customers.join(orders, on="cid", how="inner")
print(inner.shape)  # (3, 4) — cid 1 (×2), cid 2

# Left join — all left rows, nulls where no match
left = customers.join(orders, on="cid", how="left")
print(left.shape)  # (4, 4) — Charlie and Diana have null amount

# Anti join — left rows WITHOUT a match in right
no_orders = customers.join(orders, on="cid", how="anti")
print(no_orders["name"].to_list())  # ['Charlie', 'Diana']

# Join on different column names
customers.join(orders, left_on="cid", right_on="cid", suffix="_order")
python
# Asof join — match to nearest timestamp (time-series alignment)
quotes = pl.DataFrame({
    "time": [1.0, 2.0, 3.0, 4.0],
    "price": [100, 101, 102, 103],
}).cast({"time": pl.Float64})

trades = pl.DataFrame({
    "time": [1.5, 3.2],
    "qty": [50, 75],
}).cast({"time": pl.Float64})

result = trades.join_asof(quotes, on="time", strategy="backward")
print(result)
# time=1.5 matched price=100, time=3.2 matched price=102
4. Reshaping

Pivot, unpivot, explode, and transpose operations.

python
import polars as pl

# --- Pivot (long → wide) ---
long = pl.DataFrame({
    "date": ["Jan", "Jan", "Feb", "Feb"],
    "product": ["A", "B", "A", "B"],
    "sales": [100, 150, 120, 160],
})
wide = long.pivot(values="sales", index="date", columns="product")
print(wide)
# date | A   | B
# Jan  | 100 | 150
# Feb  | 120 | 160

# --- Unpivot (wide → long) ---
back_to_long = wide.unpivot(
    index="date", on=["A", "B"],
    variable_name="product", value_name="sales",
)
print(back_to_long.shape)  # (4, 3)

# --- Explode list columns ---
nested = pl.DataFrame({
    "id": [1, 2],
    "tags": [["a", "b", "c"], ["d", "e"]],
})
flat = nested.explode("tags")
print(flat.shape)  # (5, 2)
5. Data I/O

Read and write CSV, Parquet, JSON, Excel, databases, and cloud storage.

python
import polars as pl

# --- CSV ---
df = pl.read_csv("data.csv")
df.write_csv("output.csv")

# --- Parquet (recommended for performance) ---
df = pl.read_parquet("data.parquet")
df.write_parquet("output.parquet", compression="zstd")

# --- JSON / NDJSON ---
df = pl.read_ndjson("data.ndjson")
df.write_ndjson("output.ndjson")

# --- Excel ---
df = pl.read_excel("data.xlsx", sheet_name="Sheet1")
df.write_excel("output.xlsx")

# --- Lazy scan (preferred for large files) ---
lf = pl.scan_csv("large.csv")
result = lf.filter(pl.col("value") > 0).select("id", "value").collect()
print(result.shape)
python
# --- Database ---
df = pl.read_database_uri(
    "SELECT * FROM users WHERE age > 25",
    uri="postgresql://user:pass@localhost/db",
)

# --- Cloud storage (S3, GCS, Azure) ---
df = pl.read_parquet("s3://bucket/data.parquet")
df = pl.scan_parquet("gs://bucket/data/*.parquet").collect()

# --- Partitioned Parquet (Hive-style) ---
df.write_parquet("output_dir", partition_by=["year", "month"])
lf = pl.scan_parquet("output_dir/**/*.parquet")
6. Expression API

String, datetime, list, and conditional operations.

python
import polars as pl
from datetime import date

df = pl.DataFrame({
    "text": ["Hello World", "foo bar", "POLARS"],
    "dt": [date(2023, 1, 15), date(2023, 6, 30), date(2024, 12, 1)],
    "values": [[1, 2, 3], [4, 5], [6]],
})

# String operations
strings = df.select(
    lower=pl.col("text").str.to_lowercase(),
    length=pl.col("text").str.len_chars(),
    contains_o=pl.col("text").str.contains("o"),
    split=pl.col("text").str.split(" "),
)
print(strings)

# Datetime operations
dates = df.select(
    year=pl.col("dt").dt.year(),
    month=pl.col("dt").dt.month(),
    weekday=pl.col("dt").dt.weekday(),
    quarter=pl.col("dt").dt.quarter(),
)
print(dates)

# List operations
lists = df.select(
    list_len=pl.col("values").list.len(),
    list_sum=pl.col("values").list.sum(),
    first=pl.col("values").list.first(),
)
print(lists)
python
# Conditional expressions (when/then/otherwise)
df = pl.DataFrame({"score": [45, 72, 88, 95, 60]})
result = df.with_columns(
    grade=pl.when(pl.col("score") >= 90).then(pl.lit("A"))
           .when(pl.col("score") >= 80).then(pl.lit("B"))
           .when(pl.col("score") >= 70).then(pl.lit("C"))
           .otherwise(pl.lit("F")),
)
print(result)

# Null handling
df2 = pl.DataFrame({"x": [1, None, 3, None, 5]})
filled = df2.with_columns(
    filled=pl.col("x").fill_null(0),
    forward=pl.col("x").fill_null(strategy="forward"),
    is_null=pl.col("x").is_null(),
)
print(filled)

# Multi-column operations with regex selector
df3 = pl.DataFrame({"val_a": [1, 2], "val_b": [3, 4], "name": ["x", "y"]})
doubled = df3.select(pl.col("^val_.*$") * 2)
print(doubled)
7. Lazy Evaluation

Build optimized query plans before execution.

python
import polars as pl

# Lazy mode: build plan, optimize, then execute
lf = pl.scan_csv("large_dataset.csv")

result = (
    lf
    .select("user_id", "category", "amount", "date")  # projection pushdown
    .filter(pl.col("amount") > 100)                     # predicate pushdown
    .with_columns(pl.col("date").str.to_date())
    .group_by("category")
    .agg(
        pl.col("amount").sum().alias("total"),
        pl.col("user_id").n_unique().alias("unique_users"),
    )
    .sort("total", descending=True)
)

# Inspect the optimized plan
print(result.explain())

# Execute
df = result.collect()
print(df)
python
# Streaming mode for very large data
lf = pl.scan_parquet("data/*.parquet")
result = (
    lf
    .filter(pl.col("year") >= 2023)
    .group_by("region")
    .agg(pl.col("sales").sum())
    .collect(streaming=True)  # processes in batches
)
print(result)

# Sink directly to file (no full materialization)
lf.filter(pl.col("active")).sink_parquet("filtered_output.parquet")

Key Concepts

Lazy vs Eager Comparison
AspectEager (DataFrame)Lazy (LazyFrame)
Created bypl.read_*(), pl.DataFrame()pl.scan_*(), df.lazy()
ExecutionImmediateOn .collect()
OptimizationNonePredicate/projection pushdown, join reordering
StreamingNocollect(streaming=True)
Best forSmall data, interactiveLarge data, pipelines
Polars Data Types
TypePython equivalentNotes
Int8/16/32/64intChoose smallest sufficient size
UInt8/16/32/64intUnsigned
Float32/64floatFloat64 default
Booleanbool
Utf8strString type
Categorical—Low-cardinality strings (faster groupby)
Datedatetime.dateDate without time
Datetimedatetime.datetimeWith microsecond precision
Durationdatetime.timedeltaTime difference
ListlistVariable-length lists
StructdictNamed fields
NullNoneAll-null column
Key Differences from Pandas
  • No index: Row access by position only; no .loc/.iloc with labels
  • Strict typing: No silent type coercion; explicit .cast() required
  • Expressions, not methods: pl.col("x").mean() instead of df["x"].mean()
  • Parallel by default: All column operations run in parallel
  • Lazy evaluation: Available via LazyFrame for query optimization

Common Workflows

1. ETL Pipeline (CSV → Clean → Parquet)
python
import polars as pl

# Extract
lf = pl.scan_csv(
    "raw_data.csv",
    dtypes={"id": pl.Int64, "date": pl.Utf8, "amount": pl.Float64},
)

# Transform
cleaned = (
    lf
    .with_columns(pl.col("date").str.to_date("%Y-%m-%d"))
    .filter(pl.col("amount").is_not_null())
    .with_columns(
        year=pl.col("date").dt.year(),
        month=pl.col("date").dt.month(),
        amount_log=pl.col("amount").log(),
    )
    .drop_nulls()
)

# Load
cleaned.collect().write_parquet("clean_data.parquet", compression="zstd")
print("ETL complete")
2. Multi-Source Join and Aggregation
python
import polars as pl

# Simulate three data sources
users = pl.DataFrame({
    "uid": [1, 2, 3, 4],
    "name": ["Alice", "Bob", "Charlie", "Diana"],
    "region": ["East", "West", "East", "West"],
})
orders = pl.DataFrame({
    "oid": range(1, 7),
    "uid": [1, 1, 2, 3, 3, 3],
    "amount": [100, 200, 150, 50, 75, 125],
})
products = pl.DataFrame({
    "oid": range(1, 7),
    "category": ["Elec", "Books", "Elec", "Books", "Elec", "Elec"],
})

# Join → aggregate
result = (
    orders
    .join(users, on="uid", how="left")
    .join(products, on="oid", how="left")
    .group_by("region", "category")
    .agg(
        pl.col("amount").sum().alias("total"),
        pl.col("amount").mean().alias("avg_order"),
        pl.len().alias("n_orders"),
    )
    .sort("total", descending=True)
)
print(result)
3. Time-Series Feature Engineering

Uses: GroupBy, Window functions, Joins, Expression API.

  1. Load time-series data with pl.scan_csv() or pl.scan_parquet()
  2. Parse dates: .with_columns(pl.col("date").str.to_date())
  3. Sort by entity and date: .sort("entity_id", "date")
  4. Add lag features: pl.col("value").shift(n).over("entity_id")
  5. Add rolling statistics: pl.col("value").rolling_mean(window_size=7).over("entity_id")
  6. Compute percent change: (pl.col("value") - pl.col("value").shift(1)) / pl.col("value").shift(1)
  7. Collect and write: .collect().write_parquet("features.parquet")

Key Parameters

ParameterFunctionDefaultRange/OptionsEffect
how.join()"inner"inner, left, outer, cross, semi, antiJoin type
strategy.join_asof()"backward"backward, forward, nearestAsof match direction
streaming.collect()FalseTrue/FalseProcess in batches for large data
compression.write_parquet()"zstd"snappy, gzip, brotli, lz4, zstd, uncompressedParquet compression
partition_by.write_parquet()NoneList of columnsHive-style partitioning
rechunkpl.concat()FalseTrue/FalseRechunk memory after concat
aggregate_function.pivot()"first"first, sum, mean, max, min, countDuplicate handling in pivot
n_rowspl.read_csv()NonePositive intLimit rows read (for sampling)
parallelpl.read_csv()"auto"auto, columns, row_groups, noneParallel reading strategy
dtypespl.read_csv()NoneDict of column→typeOverride type inference

Best Practices

  1. Use lazy mode for large datasets: pl.scan_csv() not pl.read_csv(). Enables query optimization and streaming.

  2. Stay in the expression API: Avoid .map_elements() (runs Python, no parallelism). Prefer native Polars operations — string, datetime, list namespaces cover most needs.

  3. Select early, filter early: Place .select() and .filter() as early as possible in lazy pipelines. The optimizer can push these down but explicit placement helps.

  4. Use Categorical for low-cardinality strings: df.with_columns(pl.col("region").cast(pl.Categorical)) — dramatically speeds up groupby and joins on repeated string values.

  5. Prefer Parquet over CSV: Parquet preserves types, supports predicate pushdown, and is 5–10x smaller. Use compression="zstd" for best compression/speed balance.

  6. Anti-pattern — Python loops over rows: Never iterate rows with for row in df.iter_rows() for computation. Use expressions instead.

  7. Anti-pattern — chaining .with_columns() calls: Combine multiple column additions into a single .with_columns() call for parallel execution.

Common Recipes

Recipe: Pandas Migration Pattern
python
import polars as pl
import pandas as pd

# Convert pandas → polars
pd_df = pd.DataFrame({"col": [1, 2, 3], "group": ["a", "b", "a"]})
pl_df = pl.from_pandas(pd_df)

# Key operation mapping:
# pandas: df["col"]              → polars: df.select("col")
# pandas: df[df["col"] > 1]     → polars: df.filter(pl.col("col") > 1)
# pandas: df.assign(x=...)      → polars: df.with_columns(x=...)
# pandas: df.groupby().agg()    → polars: df.group_by().agg()
# pandas: df.groupby().transform → polars: pl.col(...).over(...)
# pandas: df.merge()            → polars: df.join()
# pandas: df.melt()             → polars: df.unpivot()

# Convert back
pd_result = pl_df.to_pandas()
Recipe: Complex Aggregation Report
python
import polars as pl

df = pl.DataFrame({
    "dept": ["Sales", "Eng", "Sales", "Eng", "Sales", "Eng"],
    "level": ["Jr", "Sr", "Sr", "Jr", "Jr", "Sr"],
    "salary": [50000, 95000, 75000, 70000, 55000, 100000],
})

report = (
    df.group_by("dept", "level")
    .agg(
        pl.col("salary").mean().alias("avg_sal"),
        pl.col("salary").median().alias("med_sal"),
        pl.col("salary").std().alias("std_sal"),
        pl.len().alias("count"),
    )
    .pivot(values="avg_sal", index="dept", columns="level")
    .with_columns(
        diff=pl.col("Sr") - pl.col("Jr"),
    )
)
print(report)
Recipe: Reading Multiple Files with Schema Alignment
python
import polars as pl
from pathlib import Path

# Read multiple CSVs with potentially different columns
files = sorted(Path("data/").glob("*.csv"))
dfs = [pl.read_csv(f) for f in files]

# Diagonal concat handles mismatched schemas (fills nulls)
combined = pl.concat(dfs, how="diagonal")
print(f"Combined: {combined.shape}")
print(f"Columns: {combined.columns}")

# Or use lazy scan for Parquet (automatic parallel)
lf = pl.scan_parquet("data/**/*.parquet")
result = lf.filter(pl.col("date") > "2023-01-01").collect()

Troubleshooting

ProblemCauseSolution
SchemaError: column not foundColumn name typo or case mismatchCheck df.columns; Polars is case-sensitive
ComputeError: cannot castType mismatch in operationUse .cast(pl.Type) explicitly
OutOfMemoryError on collectData too large for eager modeUse lf.collect(streaming=True) or filter first
Slow .map_elements()Python UDF prevents parallelismRewrite using native expressions (str/dt/list namespaces)
Join produces more rows than expectedDuplicate keys in right DataFrameDeduplicate first: df.unique(subset=["key"])
InvalidOperationError: join on different typesKey columns have different dtypesCast both to same type: .cast(pl.Int64)
.over() returns wrong valuesForgetting to include all group columnsInclude all grouping columns in .over("col1", "col2")
Parquet file unreadableWritten with incompatible compressionSpecify compression="snappy" for maximum compatibility
CSV dates read as stringsNo automatic date parsing in CSV readerParse after reading: pl.col("date").str.to_date("%Y-%m-%d")
concat fails with different schemasColumns don't match across DataFramesUse how="diagonal" to fill missing columns with null
Show full SKILL.md (590 more words)Show less

Bundled Resources

  • references/pandas_migration.md — Pandas-to-Polars migration guide with operation mapping tables (selection, filtering, column ops, aggregation, window functions, joins, reshaping, string ops, datetime ops, missing data, I/O), interoperability code, common migration patterns with side-by-side code, migration pitfalls, and migration checklist.

    • Covers: all operation mapping content from original pandas_migration.md
    • Relocated inline: key pandas differences summary → SKILL.md Key Concepts "Key Differences from Pandas" section; basic conversion recipe → SKILL.md Common Recipes "Pandas Migration Pattern"
    • Omitted: anti-pattern code examples for row iteration and sequential pipe — covered in io_best_practices.md
  • references/advanced_operations.md — Rolling windows (time-based and row-based), cumulative operations (cum_sum/max/min/prod), shift/lag/lead with grouped contexts, struct operations (create/access/unnest), list column manipulation (stats, eval, filter, explode), unique/duplicate detection, advanced sorting (nulls_last, expression-based, top-N per group), column renaming (dict, suffix/prefix/programmatic), sampling (fixed n, fraction, bootstrap), transpose, and advanced reshaping patterns (wide-long-wide, nested JSON to flat, multi-level unpivot, horizontal concat).

    • Covers: advanced operations from original operations.md + transformations from original transformations.md not in main SKILL.md
    • Relocated inline: basic selection/filtering → SKILL.md Core API section 1; groupby/aggregation → section 2; basic joins/asof → section 3; basic pivot/unpivot/explode/concat → section 4; string/date/list/conditional basics → section 6; basic window functions → section 2
    • Omitted: join performance tips (simple; covered in SKILL.md Best Practices); concatenation options (rechunk covered in Key Parameters table)
  • references/io_best_practices.md — Full I/O format guide (CSV options, Parquet options with partitioning, JSON/NDJSON, Excel multi-sheet, Arrow IPC), database connectivity (PostgreSQL, MySQL, SQLite, BigQuery), cloud storage (S3, Azure, GCS), in-memory format conversions (dict, NumPy, pandas, Arrow), format selection decision guide, schema management and error handling, expression composition and reuse patterns, column selection patterns (by type, regex, exclude), memory management (estimated_size, type optimization, streaming), pipeline functions for composable transforms, testing/debugging (query plans, schema validation, profiling), performance anti-patterns (sequential pipe, many DataFrames, in-place mutation, unspecified types), and version compatibility notes.

    • Covers: all I/O content from original io_guide.md + expression/memory/testing/performance content from original best_practices.md + format selection and version notes from original core_concepts.md
    • Relocated inline: basic CSV/Parquet/JSON/Excel/database read/write → SKILL.md Core API section 5; lazy vs eager comparison → SKILL.md Key Concepts table; basic expression context/syntax → SKILL.md section 6; parallelization/type system concepts → SKILL.md Key Concepts + Best Practices; null handling → SKILL.md section 6; categorical recommendation → SKILL.md Best Practices item 4
    • Omitted: detailed expression fundamentals (what are expressions, expression contexts) — fully covered in SKILL.md Core API; basic conditional logic examples — covered in SKILL.md section 6; basic aggregation patterns — covered in SKILL.md section 2
Per-Reference-File Disposition (Original 6 files)
Original FileLinesDispositionTarget
operations.md603ConsolidatedAdvanced ops → references/advanced_operations.md; basic selection/filter/groupby/window/string/date → SKILL.md Core API sections 1-2, 6
transformations.md550ConsolidatedReshaping/transpose → references/advanced_operations.md; basic joins/pivot/unpivot/explode/concat → SKILL.md Core API sections 3-4
io_guide.md558ConsolidatedFull I/O detail → references/io_best_practices.md; basic read/write → SKILL.md Core API section 5
best_practices.md650ConsolidatedExpression reuse, memory, testing, anti-patterns → references/io_best_practices.md; core best practices → SKILL.md Best Practices
core_concepts.md379ConsolidatedFormat selection, version notes → references/io_best_practices.md; data types, lazy/eager, parallelism → SKILL.md Key Concepts
pandas_migration.md418Migrated→ references/pandas_migration.md (expanded with window/string/datetime/missing data tables)
Intentional Omissions
  • Row iteration examples (operations.md): Not documented as a positive capability; only referenced as anti-pattern in Best Practices
  • Expression fundamentals tutorial (core_concepts.md): Expression syntax, contexts, and expansion are fully covered by SKILL.md Core API sections; a separate tutorial would duplicate
  • Detailed parallelization internals (core_concepts.md): "What gets parallelized" list omitted — users only need the Best Practices guidance to stay in the expression API
  • Copy-on-write comparison (core_concepts.md): Pandas 2.0+ copy-on-write details omitted — migration-focused, not Polars-centric
  • zarr-python — Chunked array storage; Polars can read/write Parquet that Zarr processes
  • matplotlib-scientific-plotting — Visualization; convert to pandas with .to_pandas() for plotting
  • scikit-learn-machine-learning — ML pipelines; use .to_numpy() or .to_pandas() for sklearn input

References

© jaechang-hits, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (references) in skills/scientific-computing/polars-dataframes of jaechang-hits/SciAgent-Skills.

  • SKILL.md
  • references/advanced_operations.md
  • references/io_best_practices.md
  • references/pandas_migration.md

Open the folder on GitHubat commit 82c862c

Compare with similar skills

Polars Dataframes next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Polars Dataframes compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Polars Dataframes this skilljaechang-hits/SciAgent-Skills370—~6.4kAutomated safety check: PassMIT
Verified Data Analysis with pandaspipeshub-ai/pipeshub-ai3.8k—~1.2kAutomated safety check: PassApache-2.0
Daskdavila7/claude-code-templates32k11 repos~3.5kAutomated safety check: PassMIT
DaskK-Dense-AI/scientific-agent-skills48k1 repos~4.4kAutomated safety check: NotesBSD-3-Clause
PolarsK-Dense-AI/scientific-agent-skills48k1 repos~3.3kAutomated safety check: PassMIT
Ingesting Dataancoleman/ai-design-components526—~1.9kAutomated safety check: PassMIT

Similar skills

  • Verified Data Analysis with pandas

    pipeshub-ai/pipeshub-ai

    Loads, cleans, aggregates and joins tabular data with pandas under a verification rule: every number reported must be one that the code actually printed.

    3.8k GitHub stars~1.2k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Dask

    davila7/claude-code-templates

    Parallel/distributed computing. An agent skill from davila7/claude-code-templates.

    32k GitHub starsUsed in 11 repos~3.5k tokens
    Data & AnalyticsAuto-check passed
  • Dask

    K-Dense-AI/scientific-agent-skills

    Scales pandas, NumPy, and custom Python research workflows beyond memory or across clusters with Dask.

    48k GitHub starsUsed in 1 repo~4.4k tokens
    Data & AnalyticsAuto-check: notes
  • Polars

    K-Dense-AI/scientific-agent-skills

    High-performance DataFrame library for Python ETL, analytics, and pandas migration.

    48k GitHub starsUsed in 1 repo~3.3k tokens
    Data & AnalyticsAuto-check passed
  • Ingesting Data

    ancoleman/ai-design-components

    Data ingestion patterns for loading data from cloud storage, APIs, files, and streaming sources into databases.

    526 GitHub stars~1.9k tokensUpdated 10 mo ago
    Data & AnalyticsAuto-check passed
  • Transforming Data

    ancoleman/ai-design-components

    Transform raw data into analytical assets using ETL/ELT patterns, SQL (dbt), Python (pandas/polars/PySpark), and orchestration (Airflow).

    526 GitHub stars~3k tokensUpdated 10 mo ago
    Data & AnalyticsAuto-check passed

More from jaechang-hits/SciAgent-Skills

All 163 skills in this repo
  • Molecular Visualization 3dmol

    jaechang-hits/SciAgent-Skills

    3Dmol.js WebGL molecular visualization emitted as self-contained HTML.

    370 GitHub stars~3.2k tokensUpdated 9 days ago
    Auto-check passed
  • Cobrapy Metabolic Modeling

    jaechang-hits/SciAgent-Skills

    Constraint-based (COBRA) analysis of genome-scale metabolic models: FBA, FVA, knockouts, flux sampling, production envelopes, gapfilling, media optimization.

    370 GitHub starsUsed in 1 repo~4.9k tokens
    Auto-check passed
  • Rdkit Chemdraw Cdxml

    jaechang-hits/SciAgent-Skills

    Read, write, and edit ChemDraw CDX/CDXML files with RDKit's rdkit.Chem.rdChemDraw plus direct XML editing, always paired with a rendered PNG.

    370 GitHub stars~6.9k tokensUpdated 9 days ago
    Auto-check passed
  • Pubmed Database

    jaechang-hits/SciAgent-Skills

    Programmatic PubMed access via NCBI E-utilities REST API. An agent skill from jaechang-hits/SciAgent-Skills.

    370 GitHub starsUsed in 1 repo~4.4k tokens
    Auto-check passed
  • Sciagent Skill Creator

    jaechang-hits/SciAgent-Skills

    Scaffold a new SciAgent-Skills entry. An agent skill from jaechang-hits/SciAgent-Skills.

    370 GitHub stars~2.3k tokensUpdated 9 days ago
    Auto-check passed
  • Anndata Data Structure

    jaechang-hits/SciAgent-Skills

    Annotated matrices for single-cell genomics. An agent skill from jaechang-hits/SciAgent-Skills.

    370 GitHub starsUsed in 2 repos~5.8k tokens
    Auto-check passed

Questions about Polars Dataframes

What does Polars Dataframes do?

Fast in-memory DataFrame with lazy evaluation, parallel execution, Arrow backend. Polars Dataframes is an agent skill from jaechang-hits/SciAgent-Skills. Fast in-memory DataFrame with lazy evaluation, parallel execution, Arrow backend.

When should I use Polars Dataframes?

Polars Dataframes fits situations like: tabular data in RAM (1–100 GB) when pandas is too slow; tasks that involve DataFrames.

How do I install Polars Dataframes in Claude Code?

Run `npx skills add jaechang-hits/SciAgent-Skills --skill polars-dataframes -a claude-code`. Or copy the skill folder (skills/scientific-computing/polars-dataframes in jaechang-hits/SciAgent-Skills) into .claude/skills/polars-dataframes in your project. Claude Code loads it when a task matches its description.

How do I install Polars Dataframes in Codex?

Run `npx skills add jaechang-hits/SciAgent-Skills --skill polars-dataframes -a codex`. Or copy the skill folder (skills/scientific-computing/polars-dataframes in jaechang-hits/SciAgent-Skills) into .agents/skills/polars-dataframes in your project. Codex loads it when a task matches its description.

Can I use Polars Dataframes in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jaechang-hits/SciAgent-Skills --skill polars-dataframes -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/polars-dataframes, .gemini/skills/polars-dataframes, .github/skills/polars-dataframes and .opencode/skills/polars-dataframes in your project.

What does Polars Dataframes need to run?

Going by SKILL.md and its folder, Polars Dataframes needs the command-line tools its instructions call (pip). Our summary lists: Python 3.

Does Polars Dataframes access the network?

SKILL.md names 2 domains. As links in the text: docs.pola.rs and github.com. This is read from the text; nothing was executed.

Is Polars Dataframes safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Polars Dataframes use?

Polars Dataframes is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Polars Dataframes use?

About 6.4k tokens (SKILL.md is roughly 26k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 12k tokens, read only when the agent opens those files.

What are the alternatives to Polars Dataframes?

Skills that share tags, products or a category with Polars Dataframes: Verified Data Analysis with pandas (pipeshub-ai/pipeshub-ai, 3.8k stars), Dask (davila7/claude-code-templates, 32k stars), Dask (K-Dense-AI/scientific-agent-skills, 48k stars) and Polars (K-Dense-AI/scientific-agent-skills, 48k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Polars Dataframes?

jaechang-hits (a GitHub user) maintains it in jaechang-hits/SciAgent-Skills, which has 370 GitHub stars. The repository holds 163 skills in this directory. The repository was last updated on September 29, 2026.

Source: jaechang-hits/SciAgent-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.