Agent skill

Apache Arrow

by Kilo-Org in Kilo-Org/kilo-marketplace

Expert guidance for Apache Arrow, the cross-language columnar memory format for analytics workloads.

Apache-2.0Auto-check passed

Install Apache Arrow

skills CLI
$ npx skills add Kilo-Org/kilo-marketplace --skill apache-arrow -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Kilo-Org/kilo-marketplace apache-arrow --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Kilo-Org/kilo-marketplace.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/apache-arrow .claude/skills/apache-arrow && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
apache-arrow
GitHub stars
189
Used in
1 other repo
Token cost
~2.2k tokens
SKILL.md length
279 words
Files
3
Skills in repo
87
Repo updated
First seen
Licence
Apache-2.0

At a glance

Expert guidance for Apache Arrow, the cross-language columnar memory format for analytics workloads.

  • Works in 8 steps: Parquet for storage, Arrow for compute —… → Column pruning — Always specify columns=… → Predicate pushdown — Use filters= in… → …
  • SKILL.md covers Overview, Instructions, Installation and Examples, plus 1 more section
  • Calls pip and npm

What it does

Apache Arrow is an agent skill from Kilo-Org/kilo-marketplace. Expert guidance for Apache Arrow, the cross-language columnar memory format for analytics workloads. Helps developers use Arrow for high-performance data interchange between systems, zero-copy reads, and efficient columnar processing in Python (PyArrow) and JavaScript (Arrow JS).

Its SKILL.md is about 2.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `_scores.json`). Compatibility notes: No special requirements

It works with JavaScript and Python. The repository describes itself as: Kilo Marketplace - A curated collection of Skills, MCP Servers, and Modes for enhancing AI agent capabilities across the Kilo ecosystem—including Kilo Code (VS Code extension)… The licence is Apache-2.0.

Example prompts

  • “/apache-arrow”

Requirements

  • Python 3
  • Node.js
  • Compatibility (from SKILL.md): No special requirements

Workflow steps

8 steps, taken from the first numbered list in SKILL.md.

  1. Parquet for storage, Arrow for compute — Write Parquet to disk/S3; use Arrow in-memory for processing
  2. Column pruning — Always specify columns= when reading Parquet; reading all columns wastes I/O and memory
  3. Predicate pushdown — Use filters= in Parquet reads; the reader skips row groups that don't match
  4. Zero-copy when possible — Use to_pandas(self_destruct=True) for large tables; Arrow can transfer memory ownership
  5. Batch processing for large files — Use iter_batches() instead of reading entire files into memory
  6. IPC for microservices — Arrow IPC is faster than JSON/CSV for data exchange between services
  7. Partitioned datasets for scale — Partition by date/category; queries only scan relevant partitions
  8. DuckDB for Arrow queries — DuckDB can query Arrow tables directly with zero copy: duckdb.arrow(table).query("SELECT ...")

What it can do on your machine

Read from SKILL.md and the folder at commit ff51758. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pip
    • npm

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip and npm, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    No special requirements

    From compatibility in the SKILL.md frontmatter.

Context cost

Apache Arrow loads about 2.2k tokens when it runs. Until then it costs about 73 tokens; SKILL.md has 279 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~73
When it runs · the whole SKILL.md, loaded when a task matches
~2.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Kilo-Org/kilo-marketplace at commit ff51758, republished under its Apache-2.0 licence (© Kilo-Org). 279 words, ~2,169 tokens.

Download SKILL.mdSave it as .claude/skills/apache-arrow/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
apache-arrow
description
Expert guidance for Apache Arrow, the cross-language columnar memory format for analytics workloads. Helps developers use Arrow for high-performance data interchange between systems, zero-copy reads, and efficient columnar processing in Python (PyArrow) and JavaScript (Arrow JS).
compatibility
No special requirements
metadata.author
terminal-skills
metadata.version
1.0.0
metadata.category
data
metadata.tags
data-format, columnar, interop, python, javascript

Apache Arrow — Columnar Data Format

Overview

Apache Arrow, the cross-language columnar memory format for analytics workloads. Helps developers use Arrow for high-performance data interchange between systems, zero-copy reads, and efficient columnar processing in Python (PyArrow) and JavaScript (Arrow JS).

Instructions

PyArrow — Python Interface
python
# src/data/arrow_ops.py — High-performance data operations with PyArrow
import pyarrow as pa
import pyarrow.parquet as pq
import pyarrow.compute as pc
import pyarrow.csv as pcsv

# Create Arrow tables from Python data
table = pa.table({
    "user_id": pa.array([1, 2, 3, 4, 5], type=pa.int64()),
    "name": pa.array(["Alice", "Bob", "Charlie", "Diana", "Eve"]),
    "revenue": pa.array([150.0, 320.5, 89.0, 1200.0, 45.5], type=pa.float64()),
    "signup_date": pa.array([
        "2026-01-15", "2026-01-20", "2026-02-01", "2026-02-10", "2026-03-01"
    ]).cast(pa.date32()),
    "is_active": pa.array([True, True, False, True, False]),
})

# Compute operations (vectorized, no Python loops)
high_value = pc.filter(table, pc.greater(table["revenue"], 100))
total_revenue = pc.sum(table["revenue"]).as_py()    # 1805.0
avg_revenue = pc.mean(table["revenue"]).as_py()     # 361.0
sorted_table = pc.sort_indices(table, sort_keys=[("revenue", "descending")])

# Read/write Parquet files (the standard format for Arrow data)
pq.write_table(table, "users.parquet", compression="zstd")
loaded = pq.read_table("users.parquet")

# Read with column selection and row filtering (pushdown to file)
subset = pq.read_table(
    "users.parquet",
    columns=["user_id", "revenue"],          # Only read these columns
    filters=[("revenue", ">", 100)],         # Predicate pushdown
)

# Read CSV with type inference
csv_table = pcsv.read_csv("data.csv", convert_options=pcsv.ConvertOptions(
    column_types={"amount": pa.float64(), "count": pa.int32()},
))

# Streaming reads for large files (process in batches)
parquet_file = pq.ParquetFile("large_dataset.parquet")
for batch in parquet_file.iter_batches(batch_size=10_000):
    # Process each batch (RecordBatch) without loading the full file
    filtered = pc.filter(batch, pc.greater(batch["amount"], 0))
    process_batch(filtered)
Zero-Copy Interop
python
# Arrow enables zero-copy conversion between libraries
import pyarrow as pa
import pandas as pd
import polars as pl

# Arrow → Pandas (zero-copy when possible)
arrow_table = pa.table({"x": [1, 2, 3], "y": [4.0, 5.0, 6.0]})
pandas_df = arrow_table.to_pandas()           # Near-instant for compatible types

# Pandas → Arrow
arrow_from_pandas = pa.Table.from_pandas(pandas_df)

# Arrow → Polars (zero-copy)
polars_df = pl.from_arrow(arrow_table)

# Polars → Arrow (zero-copy)
arrow_from_polars = polars_df.to_arrow()

# Arrow enables data exchange between:
# Python ↔ R (via reticulate)
# Python ↔ DuckDB (zero-copy)
# Python ↔ Spark (via PySpark)
# JavaScript ↔ WASM modules
Partitioned Datasets
python
# Work with partitioned datasets on disk or cloud storage
import pyarrow.dataset as ds

# Read a partitioned Parquet dataset (Hive-style partitioning)
# data/
#   year=2025/month=01/part-0.parquet
#   year=2025/month=02/part-0.parquet
#   year=2026/month=01/part-0.parquet

dataset = ds.dataset(
    "s3://my-bucket/events/",
    format="parquet",
    partitioning=ds.partitioning(
        pa.schema([
            ("year", pa.int32()),
            ("month", pa.int32()),
        ]),
        flavor="hive",
    ),
)

# Scan with partition pruning (only reads relevant files)
scanner = dataset.scanner(
    columns=["event_type", "user_id", "timestamp"],
    filter=(ds.field("year") == 2026) & (ds.field("month") >= 1),
)
table = scanner.to_table()

# Write partitioned dataset
ds.write_dataset(
    table,
    "output/events/",
    format="parquet",
    partitioning=ds.partitioning(
        pa.schema([("year", pa.int32()), ("month", pa.int32())]),
        flavor="hive",
    ),
    existing_data_behavior="overwrite_or_ignore",
)
Arrow IPC (Inter-Process Communication)
python
# Share data between processes without serialization overhead
import pyarrow as pa
import pyarrow.ipc as ipc

# Write Arrow IPC format (for streaming between processes)
table = pa.table({"id": [1, 2, 3], "value": [10.0, 20.0, 30.0]})

# File format (random access)
with pa.OSFile("data.arrow", "wb") as f:
    writer = ipc.new_file(f, table.schema)
    writer.write_table(table)
    writer.close()

# Stream format (append-only, lower overhead)
sink = pa.BufferOutputStream()
writer = ipc.new_stream(sink, table.schema)
writer.write_table(table)
writer.close()
buffer = sink.getvalue()    # bytes that can be sent over network/pipe

# Read back
reader = ipc.open_file("data.arrow")
loaded = reader.read_all()
JavaScript (Arrow JS)
typescript
// src/data/arrow-client.ts — Read Arrow data in the browser
import { tableFromIPC, tableToIPC } from "apache-arrow";

// Fetch Arrow IPC data from an API
async function fetchArrowData(url: string) {
  const response = await fetch(url);
  const buffer = await response.arrayBuffer();

  // Parse Arrow IPC format (zero-copy in WASM-backed implementations)
  const table = tableFromIPC(new Uint8Array(buffer));

  console.log(`Loaded ${table.numRows} rows, ${table.numCols} columns`);
  console.log("Schema:", table.schema.fields.map((f) => `${f.name}: ${f.type}`));

  // Access columns
  const ids = table.getChild("id");
  const values = table.getChild("value");

  // Iterate rows
  for (const row of table) {
    console.log(row.toJSON());  // { id: 1, value: 10.0 }
  }

  return table;
}

// Send Arrow data to a server
async function sendArrowData(url: string, table: any) {
  const buffer = tableToIPC(table);
  await fetch(url, {
    method: "POST",
    headers: { "Content-Type": "application/vnd.apache.arrow.stream" },
    body: buffer,
  });
}

Installation

bash
# Python
pip install pyarrow

# JavaScript
npm install apache-arrow

# With DuckDB (Arrow-native)
pip install duckdb    # DuckDB uses Arrow internally

Examples

Example 1: Integrating Apache Arrow into an existing application

User request:

Add Apache Arrow to my Next.js app for the AI chat feature. I want streaming responses.

The agent installs the SDK, creates an API route that initializes the Apache Arrow client, configures streaming, selects an appropriate model, and wires up the frontend to consume the stream. It handles error cases and sets up proper environment variable management for the API key.

Example 2: Optimizing zero-copy interop performance

User request:

My Apache Arrow calls are slow and expensive. Help me optimize the setup.

The agent reviews the current implementation, identifies issues (wrong model selection, missing caching, inefficient prompting, no batching), and applies optimizations specific to Apache Arrow's capabilities — adjusting model parameters, adding response caching, and implementing retry logic with exponential backoff.

Guidelines

  1. Parquet for storage, Arrow for compute — Write Parquet to disk/S3; use Arrow in-memory for processing
  2. Column pruning — Always specify columns= when reading Parquet; reading all columns wastes I/O and memory
  3. Predicate pushdown — Use filters= in Parquet reads; the reader skips row groups that don't match
  4. Zero-copy when possible — Use to_pandas(self_destruct=True) for large tables; Arrow can transfer memory ownership
  5. Batch processing for large files — Use iter_batches() instead of reading entire files into memory
  6. IPC for microservices — Arrow IPC is faster than JSON/CSV for data exchange between services
  7. Partitioned datasets for scale — Partition by date/category; queries only scan relevant partitions
  8. DuckDB for Arrow queries — DuckDB can query Arrow tables directly with zero copy: duckdb.arrow(table).query("SELECT ...")

© Kilo-Org, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in skills/apache-arrow of Kilo-Org/kilo-marketplace.

  • SKILL.md
  • LICENSE
  • _scores.json

Open the folder on GitHubat commit ff51758

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in Kilo-Org/kilo-marketplace, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Apache Arrow next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Apache Arrow compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Apache Arrow this skillKilo-Org/kilo-marketplace1891 repos~2.2kAutomated safety check: PassApache-2.0
Code Review ChecklistshareAI-lab/learn-claude-code78k5 repos~1.1kAutomated safety check: PassMIT
CCXT Crypto Exchange Library2025Emma/vibe-coding-cn23k2 repos~4.4kAutomated safety check: PassMIT
Gemini API Devgoogle-gemini/gemini-skills4.3k—~5.1kAutomated safety check: PassApache-2.0
CodeQL Security Scantrailofbits/skills7.4k—~4.6kAutomated safety check: NotesCC-BY-SA-4.0
jscpd Code Migration Trackerkucherenko/jscpd6.3k—~5kAutomated safety check: PassMIT

Similar skills

  • Code Review Checklist

    shareAI-lab/learn-claude-code

    Reviews code against a five-part checklist covering security, correctness, performance, maintainability and testing, and reports findings in a fixed format.

    78k GitHub starsUsed in 5 repos~1.1k tokens
    DevelopmentAuto-check passed
  • CCXT Crypto Exchange Library

    2025Emma/vibe-coding-cn

    Reference help for the CCXT library covering crypto exchange APIs, market data, trading and order management across 150+ exchanges in JavaScript, Python and PHP.

    23k GitHub starsUsed in 2 repos~4.4k tokens
    Business, Finance & HRAuto-check passed
  • Gemini API Dev

    google-gemini/gemini-skills

    Official

    A skill your agent uses when writing code that calls the Gemini API for text generation, multi-turn chat, multimodal understanding, image generation, video generation, speech generation (TTS), voice…

    4.3k GitHub stars~5.1k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • CodeQL Security Scan

    trailofbits/skills

    Official

    Scans a codebase for vulnerabilities with CodeQL's data flow and taint tracking in run-all or important-only modes, including data extensions for project-specific sources and sinks.

    7.4k GitHub stars~4.6k tokensUpdated 5 days ago
    SecurityAuto-check: notes
  • Measures a code port between languages or frameworks with jscpd's function-level comparison, porting tests before code and tracking what is left unmatched.

    6.3k GitHub stars~5k tokensUpdated today
    DevelopmentAuto-check passed
  • Gemini Live API Dev

    google-gemini/gemini-skills

    Official

    A skill your agent uses when building real-time, bidirectional streaming applications with the Gemini Live API, or migrating legacy Live models (2.0/2.5/3.1) to Gemini 3.8 Live.

    4.3k GitHub stars~4.6k tokensUpdated today
    Backend & APIsAuto-check passed

More from Kilo-Org/kilo-marketplace

All 87 skills in this repo
  • AzureML Project Scaffolding

    Kilo-Org/kilo-marketplace

    Sets up and maintains AzureML-ready Python projects as uv workspaces with devcontainers, a Makefile and job YAML, so local runs match cloud jobs and experiments stay reproducible.

    189 GitHub stars~3.1k tokensUpdated 8 days ago
    Auto-check: notes
  • Jupyter Notebook Builder

    Kilo-Org/kilo-marketplace

    Creates, inspects, edits and runs Jupyter notebooks, scaffolding experiment or tutorial notebooks from templates and preferring a Jupyter MCP server over raw JSON edits.

    189 GitHub stars~1.3k tokensUpdated 8 days ago
    Auto-check passed
  • Tableau Dashboard Creator

    Kilo-Org/kilo-marketplace

    Takes a plain-language dashboard request through brand setup, data exploration, planning, an interactive HTML mock and a Tableau implementation spec.

    189 GitHub stars~3.8k tokensUpdated 8 days ago
    Auto-check: notes
  • Elasticsearch File Ingest

    Kilo-Org/kilo-marketplace

    Ingest and transform data files (CSV/JSON/Parquet/Arrow IPC) into Elasticsearch with stream processing and custom transforms.

    189 GitHub stars~2.8k tokensUpdated 8 days ago
    Auto-check passed
  • Nifi Flow Layout

    Kilo-Org/kilo-marketplace

    A skill your agent uses when arranging Apache NiFi processors, process groups, ports, comments, numbering, crossing connections, dense fan-in/fan-out, or reusable readable canvas layouts.

    189 GitHub stars~1.5k tokensUpdated 8 days ago
    Auto-check passed
  • Splunk Ingest Processor Setup

    Kilo-Org/kilo-marketplace

    Render Cisco Data Fabric ingest-time routing workflows and Splunk Cloud Platform Ingest Processor setup plans with SPL2 pipelines, source types, destinations, lifecycle handoffs, queue and…

    189 GitHub stars~1.2k tokensUpdated 8 days ago
    Auto-check passed

Questions about Apache Arrow

What does Apache Arrow do?

Expert guidance for Apache Arrow, the cross-language columnar memory format for analytics workloads. Apache Arrow is an agent skill from Kilo-Org/kilo-marketplace. Expert guidance for Apache Arrow, the cross-language columnar memory format for analytics workloads.

How do I install Apache Arrow in Claude Code?

Run `npx skills add Kilo-Org/kilo-marketplace --skill apache-arrow -a claude-code`. Or copy the skill folder (skills/apache-arrow in Kilo-Org/kilo-marketplace) into .claude/skills/apache-arrow in your project. Claude Code loads it when a task matches its description.

How do I install Apache Arrow in Codex?

Run `npx skills add Kilo-Org/kilo-marketplace --skill apache-arrow -a codex`. Or copy the skill folder (skills/apache-arrow in Kilo-Org/kilo-marketplace) into .agents/skills/apache-arrow in your project. Codex loads it when a task matches its description.

Can I use Apache Arrow in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Kilo-Org/kilo-marketplace --skill apache-arrow -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/apache-arrow, .gemini/skills/apache-arrow, .github/skills/apache-arrow and .opencode/skills/apache-arrow in your project.

What does Apache Arrow need to run?

Going by SKILL.md and its folder, Apache Arrow needs the command-line tools its instructions call (pip and npm). Our summary lists: Python 3; Node.js. Compatibility (from SKILL.md): No special requirements.

Does Apache Arrow access the network?

SKILL.md contains no URLs. Its commands use pip and npm, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Apache Arrow safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Apache Arrow use?

Apache Arrow is published under the Apache-2.0 licence (from the LICENSE file in the skill folder). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Apache Arrow use?

About 2.2k tokens (SKILL.md is roughly 8.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Apache Arrow?

Skills that share tags, products or a category with Apache Arrow: Code Review Checklist (shareAI-lab/learn-claude-code, 78k stars), CCXT Crypto Exchange Library (2025Emma/vibe-coding-cn, 23k stars), Gemini API Dev (google-gemini/gemini-skills, 4.3k stars) and CodeQL Security Scan (trailofbits/skills, 7.4k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Apache Arrow?

Kilo-Org (a GitHub organization) maintains it in Kilo-Org/kilo-marketplace, which has 189 GitHub stars. The repository holds 87 skills in this directory. The repository was last updated on September 28, 2026.

Source: Kilo-Org/kilo-marketplace on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.