Agent skill

Vaex Out-of-Core DataFrames

by davila7 in davila7/claude-code-templates

Processes tabular datasets too large for RAM with Vaex: lazy DataFrames, fast aggregations, big-data plots and ML pipelines over CSV, HDF5, Arrow and Parquet.

MITAuto-check passedData & Analytics

Install Vaex Out-of-Core DataFrames

skills CLI
$ npx skills add davila7/claude-code-templates --skill vaex -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install davila7/claude-code-templates vaex --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/davila7/claude-code-templates.git skills-src && mkdir -p .claude/skills && cp -r skills-src/cli-tool/components/skills/scientific/vaex .claude/skills/vaex && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
vaex
GitHub stars
32k
Used in
12 other repos
Token cost
~1.6k tokens
SKILL.md length
540 words
Files
7 (incl. references)
Skills in repo
477
Repo updated
First seen
Licence
MIT

At a glance

Processes tabular datasets too large for RAM with Vaex: lazy DataFrames, fast aggregations, big-data plots and ML pipelines over CSV, HDF5, Arrow and Parquet.

  • Works in 6 steps: DataFrames and Data Loading → Data Processing and Manipulation → Performance and Optimization → …
  • Analyzing a CSV or HDF5 file that is larger than available memory
  • SKILL.md covers Overview, When to Use This Skill, Core Capabilities and Quick Start Pattern, plus 4 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Vaex is a Python library for lazy, out-of-core DataFrames, so datasets with billions of rows can be filtered, grouped and plotted without loading them into memory. The skill covers opening large files and converting from pandas, NumPy or Arrow, virtual columns and expressions, groupby aggregations, string and datetime handling, and missing data.

It also explains lazy evaluation and caching, batching work with `delay=True`, materializing columns when needed, and asynchronous operations, plus heatmaps, histograms and scatter plots of large data and machine learning integration. Six reference files split these areas into core DataFrames, data processing, I/O, machine learning, performance and visualization.

When your agent uses it

  • Analyzing a CSV or HDF5 file that is larger than available memory
  • Computing fast groupby statistics over a very large table
  • Plotting a heatmap or histogram of a huge dataset
  • Converting a big CSV into Parquet or Arrow

Example prompts

  • “Open trips.hdf5 with Vaex and compute the average fare by pickup hour.”
  • “Convert this 40 GB CSV to Parquet without loading it all into memory.”
  • “Make a 2D heatmap of latitude against longitude for the whole events table.”

Requirements

  • Python with `vaex` installed

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. DataFrames and Data Loading
  2. Data Processing and Manipulation
  3. Performance and Optimization
  4. Data Visualization
  5. Machine Learning Integration
  6. I/O Operations

What it can do on your machine

Read from SKILL.md and the folder at commit 14680ec. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Vaex Out-of-Core DataFrames loads about 1.6k tokens when it runs, and up to ~21k if it reads all its reference files. Until then it costs about 120 tokens; SKILL.md has 540 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~120
When it runs · the whole SKILL.md, loaded when a task matches
~1.6k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~21k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from davila7/claude-code-templates at commit 14680ec, republished under its MIT licence (© davila7). 540 words, ~1,633 tokens.

Download SKILL.mdSave it as .claude/skills/vaex/SKILL.md (or your agent's skills folder). This skill also uses 6 other files; get the full folder from GitHub.
name
vaex
description
Use this skill for processing and analyzing large tabular datasets (billions of rows) that exceed available RAM. Vaex excels at out-of-core DataFrame operations, lazy evaluation, fast aggregations, efficient visualization of big data, and machine learning on large datasets. Apply when users need to work with large CSV/HDF5/Arrow/Parquet files, perform fast statistics on massive datasets, create visualizations of big data, or build ML pipelines that don't fit in memory.

Vaex

Overview

Vaex is a high-performance Python library designed for lazy, out-of-core DataFrames to process and visualize tabular datasets that are too large to fit into RAM. Vaex can process over a billion rows per second, enabling interactive data exploration and analysis on datasets with billions of rows.

When to Use This Skill

Use Vaex when:

  • Processing tabular datasets larger than available RAM (gigabytes to terabytes)
  • Performing fast statistical aggregations on massive datasets
  • Creating visualizations and heatmaps of large datasets
  • Building machine learning pipelines on big data
  • Converting between data formats (CSV, HDF5, Arrow, Parquet)
  • Needing lazy evaluation and virtual columns to avoid memory overhead
  • Working with astronomical data, financial time series, or other large-scale scientific datasets

Core Capabilities

Vaex provides six primary capability areas, each documented in detail in the references directory:

1. DataFrames and Data Loading

Load and create Vaex DataFrames from various sources including files (HDF5, CSV, Arrow, Parquet), pandas DataFrames, NumPy arrays, and dictionaries. Reference references/core_dataframes.md for:

  • Opening large files efficiently
  • Converting from pandas/NumPy/Arrow
  • Working with example datasets
  • Understanding DataFrame structure
2. Data Processing and Manipulation

Perform filtering, create virtual columns, use expressions, and aggregate data without loading everything into memory. Reference references/data_processing.md for:

  • Filtering and selections
  • Virtual columns and expressions
  • Groupby operations and aggregations
  • String operations and datetime handling
  • Working with missing data
3. Performance and Optimization

Leverage Vaex's lazy evaluation, caching strategies, and memory-efficient operations. Reference references/performance.md for:

  • Understanding lazy evaluation
  • Using delay=True for batching operations
  • Materializing columns when needed
  • Caching strategies
  • Asynchronous operations
4. Data Visualization

Create interactive visualizations of large datasets including heatmaps, histograms, and scatter plots. Reference references/visualization.md for:

  • Creating 1D and 2D plots
  • Heatmap visualizations
  • Working with selections
  • Customizing plots and subplots
5. Machine Learning Integration

Build ML pipelines with transformers, encoders, and integration with scikit-learn, XGBoost, and other frameworks. Reference references/machine_learning.md for:

  • Feature scaling and encoding
  • PCA and dimensionality reduction
  • K-means clustering
  • Integration with scikit-learn/XGBoost/CatBoost
  • Model serialization and deployment
Show full SKILL.md (220 more words)Show less
6. I/O Operations

Efficiently read and write data in various formats with optimal performance. Reference references/io_operations.md for:

  • File format recommendations
  • Export strategies
  • Working with Apache Arrow
  • CSV handling for large files
  • Server and remote data access

Quick Start Pattern

For most Vaex tasks, follow this pattern:

python
import vaex

# 1. Open or create DataFrame
df = vaex.open('large_file.hdf5')  # or .csv, .arrow, .parquet
# OR
df = vaex.from_pandas(pandas_df)

# 2. Explore the data
print(df)  # Shows first/last rows and column info
df.describe()  # Statistical summary

# 3. Create virtual columns (no memory overhead)
df['new_column'] = df.x ** 2 + df.y

# 4. Filter with selections
df_filtered = df[df.age > 25]

# 5. Compute statistics (fast, lazy evaluation)
mean_val = df.x.mean()
stats = df.groupby('category').agg({'value': 'sum'})

# 6. Visualize
df.plot1d(df.x, limits=[0, 100])
df.plot(df.x, df.y, limits='99.7%')

# 7. Export if needed
df.export_hdf5('output.hdf5')

Working with References

The reference files contain detailed information about each capability area. Load references into context based on the specific task:

  • Basic operations: Start with references/core_dataframes.md and references/data_processing.md
  • Performance issues: Check references/performance.md
  • Visualization tasks: Use references/visualization.md
  • ML pipelines: Reference references/machine_learning.md
  • File I/O: Consult references/io_operations.md

Best Practices

  1. Use HDF5 or Apache Arrow formats for optimal performance with large datasets
  2. Leverage virtual columns instead of materializing data to save memory
  3. Batch operations using delay=True when performing multiple calculations
  4. Export to efficient formats rather than keeping data in CSV
  5. Use expressions for complex calculations without intermediate storage
  6. Profile with df.stat() to understand memory usage and optimize operations

Common Patterns

Pattern: Converting Large CSV to HDF5
python
import vaex

# Open large CSV (processes in chunks automatically)
df = vaex.from_csv('large_file.csv')

# Export to HDF5 for faster future access
df.export_hdf5('large_file.hdf5')

# Future loads are instant
df = vaex.open('large_file.hdf5')
Pattern: Efficient Aggregations
python
# Use delay=True to batch multiple operations
mean_x = df.x.mean(delay=True)
std_y = df.y.std(delay=True)
sum_z = df.z.sum(delay=True)

# Execute all at once
results = vaex.execute([mean_x, std_y, sum_z])
Pattern: Virtual Columns for Feature Engineering
python
# No memory overhead - computed on the fly
df['age_squared'] = df.age ** 2
df['full_name'] = df.first_name + ' ' + df.last_name
df['is_adult'] = df.age >= 18

Resources

This skill includes reference documentation in the references/ directory:

  • core_dataframes.md - DataFrame creation, loading, and basic structure
  • data_processing.md - Filtering, expressions, aggregations, and transformations
  • performance.md - Optimization strategies and lazy evaluation
  • visualization.md - Plotting and interactive visualizations
  • machine_learning.md - ML pipelines and model integration
  • io_operations.md - File formats and data import/export

© davila7, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 6 other files (references) in cli-tool/components/skills/scientific/vaex of davila7/claude-code-templates.

  • SKILL.md
  • references/core_dataframes.md
  • references/data_processing.md
  • references/io_operations.md
  • references/machine_learning.md
  • references/performance.md
  • references/visualization.md

Open the folder on GitHubat commit 14680ec

Used in 12 other repositories

We found 14 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 12 other GitHub owners. This page covers the copy in davila7/claude-code-templates, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Vaex Out-of-Core DataFrames next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Vaex Out-of-Core DataFrames compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Vaex Out-of-Core DataFrames this skilldavila7/claude-code-templates32k12 repos~1.6kAutomated safety check: PassMIT
Python Executorcortega26/chile-hub1132 repos~1.5kAutomated safety check: PassMIT
Hybrid-Engine Data Analysiscode-yeongyu/oh-my-openagent70k—~1.4kAutomated safety check: PassCustom licence
CSV Data Summarizercoffeefuelbump/csv-data-summarizer-claude-skill4682 repos~1.4kAutomated safety check: PassNone
Verified Data Analysis with pandaspipeshub-ai/pipeshub-ai3.8k—~1.2kAutomated safety check: PassApache-2.0
Analytics Data AnalysisMindrally/skills268—~1.6kAutomated safety check: PassApache-2.0

Similar skills

  • Python Executor

    cortega26/chile-hub

    Execute Python code in a safe sandboxed environment via [inference.sh](https://inference.sh).

    113 GitHub starsUsed in 2 repos~1.5k tokens
    Data & AnalyticsAuto-check passed
  • Hybrid-Engine Data Analysis

    code-yeongyu/oh-my-openagent

    Analyzes CSV, Parquet and JSON data with DuckDB, Polars, numpy and matplotlib, preferring a persistent kernel over repeated one-shot processes.

    70k GitHub stars~1.4k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • CSV Data Summarizer

    coffeefuelbump/csv-data-summarizer-claude-skill

    Analyzes CSV files, generates summary stats, and plots quick visualizations using Python and pandas.

    468 GitHub starsUsed in 2 repos~1.4k tokens
    Data & AnalyticsAuto-check passed
  • Verified Data Analysis with pandas

    pipeshub-ai/pipeshub-ai

    Loads, cleans, aggregates and joins tabular data with pandas under a verification rule: every number reported must be one that the code actually printed.

    3.8k GitHub stars~1.2k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Analytics Data Analysis

    Mindrally/skills

    Best practices for analytics, data analysis, and visualization using Python, pandas, matplotlib, seaborn, and Jupyter notebooks.

    268 GitHub stars~1.6k tokensUpdated 1 mo ago
    Data & AnalyticsAuto-check passed
  • Dask

    K-Dense-AI/scientific-agent-skills

    Scales pandas, NumPy, and custom Python research workflows beyond memory or across clusters with Dask.

    48k GitHub starsUsed in 1 repo~4.4k tokens
    Data & AnalyticsAuto-check: notes

More from davila7/claude-code-templates

All 477 skills in this repo
  • Perplexity Web Search

    davila7/claude-code-templates

    Runs web-grounded searches through Perplexity's Sonar models over OpenRouter for current events, recent literature and cited facts beyond the model's training cutoff.

    32k GitHub starsUsed in 12 repos~3.5k tokens
    Auto-check: notes
  • Neuropixels Data Analysis

    davila7/claude-code-templates

    Analyzes Neuropixels recordings from SpikeGLX or Open Ephys through preprocessing, drift correction, Kilosort4 spike sorting, quality metrics and curation.

    32k GitHub starsUsed in 10 repos~2.8k tokens
    Auto-check passed
  • Scientific Venue Templates

    davila7/claude-code-templates

    Supplies LaTeX templates and formatting rules for journals, conferences, posters, and grant proposals, then can check a draft against them.

    32k GitHub starsUsed in 9 repos~5.1k tokens
    Auto-check: notes
  • Brand Voice Content Creator

    davila7/claude-code-templates

    Analyzes a brand's existing writing to lock in a consistent voice, then builds SEO blog posts and platform-specific social content around it.

    32k GitHub starsUsed in 3 repos~1.9k tokens
    Auto-check passed
  • CAPA Officer

    davila7/claude-code-templates

    Guides corrective and preventive action (CAPA) work in a quality management system, from initiation and root cause analysis through effectiveness verification.

    32k GitHub starsUsed in 1 repo~2k tokens
    Auto-check passed
  • Fda Consultant Specialist

    davila7/claude-code-templates

    Senior FDA consultant and specialist for medical device companies including HIPAA compliance and requirement management.

    32k GitHub starsUsed in 1 repo~2.7k tokens
    Auto-check passed

Questions about Vaex Out-of-Core DataFrames

What does Vaex Out-of-Core DataFrames do?

Processes tabular datasets too large for RAM with Vaex: lazy DataFrames, fast aggregations, big-data plots and ML pipelines over CSV, HDF5, Arrow and Parquet. Vaex is a Python library for lazy, out-of-core DataFrames, so datasets with billions of rows can be filtered, grouped and plotted without loading them into memory. The skill covers opening large files and converting from pandas, NumPy or Arrow, virtual columns and expressions, groupby aggregations, string and datetime handling, and missing data.

When should I use Vaex Out-of-Core DataFrames?

Vaex Out-of-Core DataFrames fits situations like: analyzing a CSV or HDF5 file that is larger than available memory; computing fast groupby statistics over a very large table; plotting a heatmap or histogram of a huge dataset; converting a big CSV into Parquet or Arrow.

How do I install Vaex Out-of-Core DataFrames in Claude Code?

Run `npx skills add davila7/claude-code-templates --skill vaex -a claude-code`. Or copy the skill folder (cli-tool/components/skills/scientific/vaex in davila7/claude-code-templates) into .claude/skills/vaex in your project. Claude Code loads it when a task matches its description.

How do I install Vaex Out-of-Core DataFrames in Codex?

Run `npx skills add davila7/claude-code-templates --skill vaex -a codex`. Or copy the skill folder (cli-tool/components/skills/scientific/vaex in davila7/claude-code-templates) into .agents/skills/vaex in your project. Codex loads it when a task matches its description.

Can I use Vaex Out-of-Core DataFrames in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add davila7/claude-code-templates --skill vaex -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/vaex, .gemini/skills/vaex, .github/skills/vaex and .opencode/skills/vaex in your project.

What does Vaex Out-of-Core DataFrames need to run?

SKILL.md names no scripts, command-line tools or credentials: Vaex Out-of-Core DataFrames is instructions for the agent only. Our summary lists: Python with `vaex` installed.

Does Vaex Out-of-Core DataFrames access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Vaex Out-of-Core DataFrames safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Vaex Out-of-Core DataFrames use?

Vaex Out-of-Core DataFrames is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Vaex Out-of-Core DataFrames use?

About 1.6k tokens (SKILL.md is roughly 6.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 20k tokens, read only when the agent opens those files.

What are the alternatives to Vaex Out-of-Core DataFrames?

Skills that share tags, products or a category with Vaex Out-of-Core DataFrames: Python Executor (cortega26/chile-hub, 113 stars), Hybrid-Engine Data Analysis (code-yeongyu/oh-my-openagent, 70k stars), CSV Data Summarizer (coffeefuelbump/csv-data-summarizer-claude-skill, 468 stars) and Verified Data Analysis with pandas (pipeshub-ai/pipeshub-ai, 3.8k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Vaex Out-of-Core DataFrames?

davila7 (a GitHub user) maintains it in davila7/claude-code-templates, which has 32,463 GitHub stars. The repository holds 477 skills in this directory. The repository was last updated on October 8, 2026.

Source: davila7/claude-code-templates on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.