Official agent skill

Accelerated Computing Cudf

by NVIDIA in NVIDIA/skills

Official NVIDIA-authored guidance for NVIDIA cuDF GPU DataFrames, pandas acceleration, dask-cuDF, ETL, joins, groupby, CSV/Parquet I/O, nullable semantics, and multi-GPU DataFrame workloads.

OfficialApache-2.0Auto-check passedData & Analytics

Install Accelerated Computing Cudf

skills CLI
$ npx skills add NVIDIA/skills --skill accelerated-computing-cudf -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NVIDIA/skills accelerated-computing-cudf --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/accelerated-computing-cudf .claude/skills/accelerated-computing-cudf && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
accelerated-computing-cudf
GitHub stars
3.5k
Token cost
~2.3k tokens
SKILL.md length
982 words
Files
53 (incl. references)
Skills in repo
380
Repo updated
First seen
Licence
Apache-2.0

At a glance

Official NVIDIA-authored guidance for NVIDIA cuDF GPU DataFrames, pandas acceleration, dask-cuDF, ETL, joins, groupby, CSV/Parquet I/O, nullable semantics, and multi-GPU DataFrame workloads.

  • Works in 6 steps: Choose the right cuDF path. Use… → Size gate: 100K rows minimum. Below… → Keep conversions at boundaries. Use… → …
  • Tasks that involve DataFrames
  • SKILL.md covers Compatibility, Naming, Role and Critical Rules, plus 6 more sections
  • Runs Python scripts from its folder; calls python

What it does

Accelerated Computing Cudf is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Official NVIDIA-authored guidance for NVIDIA cuDF GPU DataFrames, pandas acceleration, dask-cuDF, ETL, joins, groupby, CSV/Parquet I/O, nullable semantics, and multi-GPU DataFrame workloads.

Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 62 other files, including reference files (for example `BENCHMARK.md`, `evals/evals.json` and `evals/files/cudf-apply-udf/code/generate_data.py`).

It sits in Data & Analytics, covering DataFrames. It works with NVIDIA AI Platform, pandas, Dask and CUDA. The repository describes itself as: Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve DataFrames

Example prompts

  • “/accelerated-computing-cudf”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Choose the right cuDF path. Use cudf.pandas for broad compatibility or minimal-change acceleration. Use explicit cuDF when the user asks…
  2. Size gate: 100K rows minimum. Below that, GPU transfer overhead usually beats the speedup; use small data for correctness and benchmark…
  3. Keep conversions at boundaries. Use .to_pandas(), .values, or .numpy() for display, plotting, CPU-only libraries, or final output…
  4. Float32 is your friend. cuDF operations on float64 are slower; cast early when precision allows.
  5. Validate semantics on representative slices. For null handling, joins, time series, reshape, or grouped logic, keep a small pandas…
  6. For data > GPU memory, move to dask-cuDF with enable_cudf_spill=True. See references/dask-cudf-patterns.md.

What it can do on your machine

Read from SKILL.md and the folder at commit 0e0d506. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python, from the files we listed), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • docs.nvidia.com
    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Accelerated Computing Cudf loads about 2.3k tokens when it runs, and up to ~6.8k if it reads all its reference files. Until then it costs about 54 tokens; SKILL.md has 982 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~54
When it runs · the whole SKILL.md, loaded when a task matches
~2.3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~6.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from NVIDIA/skills at commit 0e0d506, republished under its Apache-2.0 licence (© NVIDIA). 982 words, ~2,341 tokens.

Download SKILL.mdSave it as .claude/skills/accelerated-computing-cudf/SKILL.md (or your agent's skills folder). This skill also uses 52 other files; get the full folder from GitHub.
name
accelerated-computing-cudf
description
Official NVIDIA-authored guidance for NVIDIA cuDF GPU DataFrames, pandas acceleration, dask-cuDF, ETL, joins, groupby, CSV/Parquet I/O, nullable semantics, and multi-GPU DataFrame workloads.
license
CC-BY-4.0 AND Apache-2.0
metadata.author
NVIDIA
metadata.tags
cudf, dataframes, pandas, dask-cudf, etl

cuDF & dask-cuDF Implementer's Guide

Compatibility

  • Release tracked by this skill: 26.04.
  • Requires NVIDIA Volta or newer on CUDA 12, or Turing or newer on CUDA 13. Release 26.04 supports CUDA 12.2-12.9 with driver 535+ or CUDA 13.0-13.1 with driver 580+, and Python 3.11-3.14. cuDF sweet spot: >100K rows.

Naming

Use NVIDIA library-first wording in user-facing answers. Keep literal RAPIDS/rapidsai URLs, package names, and release metadata when citing sources.

Role

You are a cuDF expert helping an implementer work with GPU DataFrames. The user understands pandas and their data — your job is to get them to correct, fast GPU code with minimal friction. Choose the path from the user's intent: cudf.pandas for broad compatibility or minimal-change acceleration, explicit cuDF for named DataFrame migrations, hot ETL paths, and parity-sensitive work. Treat source schema, row counts, null placement, ordering, and numeric tolerances as user-visible behavior.

Critical Rules

  1. Choose the right cuDF path. Use cudf.pandas for broad compatibility or minimal-change acceleration. Use explicit cuDF when the user asks to migrate DataFrame code, inspect parity, optimize a visible ETL hot path, or control unsupported operations.
  2. Size gate: 100K rows minimum. Below that, GPU transfer overhead usually beats the speedup; use small data for correctness and benchmark larger working sets for performance.
  3. Keep conversions at boundaries. Use .to_pandas(), .values, or .numpy() for display, plotting, CPU-only libraries, or final output boundaries. Keep intermediate ETL data on GPU.
  4. Float32 is your friend. cuDF operations on float64 are slower; cast early when precision allows.
  5. Validate semantics on representative slices. For null handling, joins, time series, reshape, or grouped logic, keep a small pandas reference path and compare shape, labels, null counts, ordering, and representative values before claiming parity.
  6. For data > GPU memory, move to dask-cuDF with enable_cudf_spill=True. See references/dask-cudf-patterns.md.

Three Paths to GPU DataFrames

Path 1: cudf.pandas Accelerator (Compatibility / Minimal Change)

Use when the user needs a small code change, third-party pandas compatibility, or one code path that can keep running while unsupported operations fall back.

Jupyter/IPython:

python
%load_ext cudf.pandas
import pandas as pd   # now GPU-backed; falls back silently for unsupported ops

Script:

bash
python -m cudf.pandas my_script.py

With multiprocessing:

python
import cudf.pandas
cudf.pandas.install()   # must come BEFORE pandas import, before Pool creation
from multiprocessing import Pool

Confirm acceleration with the cudf.pandas profiler before claiming speedup. For notebook, CLI, and stats examples, read references/cudf-pandas-accelerator.md. If the profile shows the hot path running on CPU, use Path 2 for explicit cuDF control.

Path 2: Explicit cuDF API

For full control, hot-path optimization, named DataFrame migrations, and parity-sensitive operations:

python
import cudf

# Read data directly to GPU
df = cudf.read_parquet("data.parquet")

# Operations mirror pandas
result = df.groupby("key")["value"].sum()
merged = df.merge(lookup, on="id", how="left")
filtered = df[df["amount"] > 1000]

# String operations
df["clean"] = df["name"].str.strip().str.lower()

# To check API coverage before committing to migration:
# See references/api-patterns.md for known gaps and workarounds

Keep data on GPU end-to-end. Only call .to_pandas() at the very end for display or CPU or non-GPU handoff.

Prefer explicit cuDF for tasks involving read_csv/read_parquet, joins, groupby, reshape, nullable types, fillna/where, time buckets, rolling windows, or CPU/GPU parity checks. Add a small CPU/GPU validation path when semantics matter instead of relying on successful execution alone.

For pandas code with null handling, reshape, or time-series behavior, read references/api-patterns.md for the relevant semantic checklist before rewriting. A cudf.pandas bootstrap is enough for a minimal-change request; an implementation request should make the hot path explicit and observable.

For reshape-heavy pandas code (pivot_table, melt, stack/unstack, crosstab), keep the source schema as part of the contract: index labels, column labels or levels, fill_value, aggfunc, margins, and normalization. Use explicit cuDF where the equivalent is supported; use cudf.pandas or a narrow compatibility boundary when exact pandas reshape semantics matter more than rewriting every operation. Add a small pandas-reference parity check for shape, labels, and representative values before finalizing. See references/api-patterns.md.

Path 3: dask-cuDF (Multi-GPU / Large Data)

When dataset exceeds GPU memory. See references/dask-cudf-patterns.md for full patterns.

python
from dask_cuda import LocalCUDACluster
from dask.distributed import Client
import dask_cudf

cluster = LocalCUDACluster(enable_cudf_spill=True)  # one worker per GPU
client = Client(cluster)

ddf = dask_cudf.read_parquet("s3://bucket/data/*.parquet")
result = ddf.groupby("key").agg({"value": "sum"}).compute()
Show full SKILL.md (414 more words)Show less

Memory Management

Enable spill before OOM happens (not after):

python
import cudf
cudf.set_option("spill", True)   # spill to host RAM when GPU is full

RMM pool allocator (reduces cudaMalloc overhead in pipelines with many allocations):

python
import rmm
rmm.set_current_device_resource(rmm.mr.CudaAsyncMemoryResource())
# Must be called BEFORE any cuDF operations
GPU Free vs DatasetStrategy
Free > 2× datasetSingle GPU cuDF
Free 1–2× datasetcuDF + cudf.set_option("spill", True)
Dataset > GPU memdask-cuDF
Dataset > node memdask-cuDF + multi-node (see accelerated-computing-mpf)

Troubleshooting

No speedup vs pandas:

  • Data < 100K rows? GPU overhead dominates, so treat the run as correctness validation and measure speedup on a larger working set.
  • Run %%cudf.pandas.profile — high CPU % means many fallbacks. Identify and fix those ops.
  • Check references/api-patterns.md for known gaps.

OOM (CUDA out of memory):

  1. Enable spill: cudf.set_option("spill", True)
  2. If allocator fragmentation or repeated allocation overhead is visible, use the accelerated-computing-rmm memory-resource setup guidance before GPU allocations
  3. Still failing: move to dask-cuDF

AttributeError / NotImplementedError:

  • Check references/api-patterns.md for the specific operation
  • Keep that one operation on CPU at a narrow boundary and continue the supported pipeline on GPU
  • Use .to_pandas() only for the unsupported op, then .from_pandas() back

Wrong results vs pandas:

  • Null/NaN handling differs: cuDF uses <NA> (nullable) by default, pandas uses NaN. See references/api-patterns.md.
  • Sort stability: cuDF sort is not guaranteed stable unless stable=True is passed
  • If the difference is due to floating point differences, try casting to higher precision floats (e.g. float64 instead of float32). If the results are still different, stop. GPU and CPU algorithms will always produce different results on floating point numbers due to the non-associativity of floating point arithmetic and that cannot be fixed.

Nullable and Fill Semantics

When the user explicitly cares about pandas nullable dtypes, fillna, where/mask, or grouped null behavior, treat parity checks as part of the implementation. See references/api-patterns.md for nullable dtype examples.

  • Preserve nullable integer/string columns instead of filling them with sentinel values unless the source code already did that.
  • Keep where/mask semantics when they encode a condition. Use broad fillna only when the condition is exactly null-only.
  • Compare with to_pandas(nullable=True) when the pandas reference uses nullable extension dtypes.
  • Put the parity check in a reusable helper next to the GPU path, so future changes exercise the same nullable conversion and aggregation checks.
  • Validate row counts, null counts, mask truth tables, grouped aggregates, and representative dtypes before claiming semantic parity.

Reference Files

  • references/cudf-pandas-accelerator.md — Profiling, fallback detection, cudf.pandas deep dive
  • references/api-patterns.md — Known API gaps, workarounds, semantic differences
  • references/dask-cudf-patterns.md — Multi-GPU patterns, best practices, partition tuning

External Documentation

Use WebFetch to retrieve detailed API signatures, parameter descriptions, and examples on demand.

© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 52 other files (references) in skills/accelerated-computing-cudf of NVIDIA/skills.

  • SKILL.md
  • BENCHMARK.md
  • evals/evals.json
  • evals/files/cudf-apply-udf/code/generate_data.py
  • evals/files/cudf-apply-udf/code/udf_pipeline.py
  • evals/files/cudf-csv-etl/code/etl_pipeline.py
  • evals/files/cudf-csv-etl/code/generate_data.py
  • evals/files/cudf-groupby-agg/code/generate_data.py
  • evals/files/cudf-groupby-agg/code/groupby_analysis.py
  • evals/files/cudf-multi-join/code/generate_data.py
  • evals/files/cudf-multi-join/code/multi_join.py
  • … and 42 more

Open the folder on GitHubat commit 0e0d506

Compare with similar skills

Accelerated Computing Cudf next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Accelerated Computing Cudf compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Accelerated Computing Cudf this skillNVIDIA/skills3.5k—~2.3kAutomated safety check: PassApache-2.0
Daskdavila7/claude-code-templates32k11 repos~3.5kAutomated safety check: PassMIT
DaskK-Dense-AI/scientific-agent-skills48k1 repos~4.4kAutomated safety check: NotesBSD-3-Clause
Optimize For GPUK-Dense-AI/scientific-agent-skills48k1 repos~3.4kAutomated safety check: PassMIT
Optimize For GPUmajiayu000/claude-skill-registry6661 repos~8.5kAutomated safety check: PassMIT
Dataframe WorkflowsVectorSpaceLab/AREX-Skill328—~1kAutomated safety check: PassBSD-3-Clause

Similar skills

  • Dask

    davila7/claude-code-templates

    Parallel/distributed computing. An agent skill from davila7/claude-code-templates.

    32k GitHub starsUsed in 11 repos~3.5k tokens
    Data & AnalyticsAuto-check passed
  • Dask

    K-Dense-AI/scientific-agent-skills

    Scales pandas, NumPy, and custom Python research workflows beyond memory or across clusters with Dask.

    48k GitHub starsUsed in 1 repo~4.4k tokens
    Data & AnalyticsAuto-check: notes
  • Optimize For GPU

    K-Dense-AI/scientific-agent-skills

    GPU-accelerates scientific Python on NVIDIA hardware and verifies that the result is correct and faster.

    48k GitHub starsUsed in 1 repo~3.4k tokens
    Data & AnalyticsAuto-check passed
  • Optimize For GPU

    majiayu000/claude-skill-registry

    GPU-accelerate Python code using CuPy, Numba CUDA, Warp, cuDF, cuML, cuGraph, KvikIO, cuCIM, cuxfilter, cuVS, cuSpatial, and RAFT.

    666 GitHub starsUsed in 1 repo~8.5k tokens
    Data & AnalyticsAuto-check passed
  • Dataframe Workflows

    VectorSpaceLab/AREX-Skill

    Use this Dask sub-skill for Dask DataFrame creation, CSV/Parquet/JSON/SQL IO, partitions and divisions, groupby/aggregation, joins/merge, shuffle, repartitioning, categorical/string/pyarrow…

    328 GitHub stars~1k tokensUpdated 1 mo ago
    Data & AnalyticsAuto-check passed
  • Polars Dataframes

    jaechang-hits/SciAgent-Skills

    Fast in-memory DataFrame with lazy evaluation, parallel execution, Arrow backend.

    370 GitHub stars~6.4k tokensUpdated 8 days ago
    Data & AnalyticsAuto-check passed

More from NVIDIA/skills

All 380 skills in this repo
  • Official

    A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.

    3.5k GitHub starsUsed in 1 repo~4.5k tokens
    Auto-check passed
  • Official

    Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.

    3.5k GitHub stars~2.9k tokensUpdated today
    Auto-check passed
  • Official

    Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.

    3.5k GitHub stars~4.8k tokensUpdated today
    Auto-check passed
  • Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.

    3.5k GitHub stars~5k tokensUpdated today
    Auto-check: notes
  • Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.

    3.5k GitHub stars~4.7k tokensUpdated today
    Auto-check: notes
  • Official

    Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.

    3.5k GitHub stars~2.7k tokensUpdated today
    Auto-check: notes

Questions about Accelerated Computing Cudf

What does Accelerated Computing Cudf do?

Official NVIDIA-authored guidance for NVIDIA cuDF GPU DataFrames, pandas acceleration, dask-cuDF, ETL, joins, groupby, CSV/Parquet I/O, nullable semantics, and multi-GPU DataFrame workloads. Accelerated Computing Cudf is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Official NVIDIA-authored guidance for NVIDIA cuDF GPU DataFrames, pandas acceleration, dask-cuDF, ETL, joins, groupby, CSV/Parquet I/O, nullable semantics, and multi-GPU DataFrame workloads.

When should I use Accelerated Computing Cudf?

Accelerated Computing Cudf fits situations like: tasks that involve DataFrames.

How do I install Accelerated Computing Cudf in Claude Code?

Run `npx skills add NVIDIA/skills --skill accelerated-computing-cudf -a claude-code`. Or copy the skill folder (skills/accelerated-computing-cudf in NVIDIA/skills) into .claude/skills/accelerated-computing-cudf in your project. Claude Code loads it when a task matches its description.

How do I install Accelerated Computing Cudf in Codex?

Run `npx skills add NVIDIA/skills --skill accelerated-computing-cudf -a codex`. Or copy the skill folder (skills/accelerated-computing-cudf in NVIDIA/skills) into .agents/skills/accelerated-computing-cudf in your project. Codex loads it when a task matches its description.

Can I use Accelerated Computing Cudf in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill accelerated-computing-cudf -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/accelerated-computing-cudf, .gemini/skills/accelerated-computing-cudf, .github/skills/accelerated-computing-cudf and .opencode/skills/accelerated-computing-cudf in your project.

What does Accelerated Computing Cudf need to run?

Going by SKILL.md and its folder, Accelerated Computing Cudf needs Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python 3.

Does Accelerated Computing Cudf access the network?

SKILL.md names 2 domains. As links in the text: docs.nvidia.com and github.com. This is read from the text; nothing was executed.

Is Accelerated Computing Cudf safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Accelerated Computing Cudf use?

Accelerated Computing Cudf is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Accelerated Computing Cudf use?

About 2.3k tokens (SKILL.md is roughly 9.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4.5k tokens, read only when the agent opens those files.

What are the alternatives to Accelerated Computing Cudf?

Skills that share tags, products or a category with Accelerated Computing Cudf: Dask (davila7/claude-code-templates, 32k stars), Dask (K-Dense-AI/scientific-agent-skills, 48k stars), Optimize For GPU (K-Dense-AI/scientific-agent-skills, 48k stars) and Optimize For GPU (majiayu000/claude-skill-registry, 666 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Accelerated Computing Cudf?

NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,534 GitHub stars. The repository holds 380 skills in this directory. The repository was last updated on October 7, 2026.

Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.