Agent skill

Io Bound Data Processing

by pproenca in pproenca/dot-skills

Processing, transforming, or moving datasets that may exceed RAM on a single low-compute box — covers memory discipline (streaming, generators, dtype shrinkage), I/O access patterns (sequential vs…

MITAuto-check passedData & Analytics

Install Io Bound Data Processing

skills CLI
$ npx skills add pproenca/dot-skills --skill io-bound-data-processing -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install pproenca/dot-skills io-bound-data-processing --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/pproenca/dot-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/.experimental/io-bound-data-processing .claude/skills/io-bound-data-processing && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
io-bound-data-processing
GitHub stars
214
Token cost
~3.3k tokens
SKILL.md length
933 words
Files
47 (incl. references, assets)
Skills in repo
182
Repo updated
First seen
Licence
MIT

At a glance

Processing, transforming, or moving datasets that may exceed RAM on a single low-compute box — covers memory discipline (streaming, generators, dtype shrinkage), I/O access patterns (sequential vs…

  • Works in 9 steps: Memory Discipline (CRITICAL) → I/O Access Patterns (CRITICAL) → Data Format & Encoding (HIGH) → …
  • Process a large file
  • SKILL.md covers When to Apply, Rule Categories By Priority, Quick Reference and How to Use, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Io Bound Data Processing is an agent skill from pproenca/dot-skills. Processing, transforming, or moving datasets that may exceed RAM on a single low-compute box — covers memory discipline (streaming, generators, dtype shrinkage), I/O access patterns (sequential vs random, mmap, async), data formats (Parquet vs CSV vs JSON, predicate pushdown), chunking & batching, spill-to-disk (external merge sort, DuckDB/Polars), pipelining (bounded queues, backpressure, checkpointing), codec selection (zstd/lz4/gzip), concurrency for I/O-bound workloads (asyncio, threads, prefetch), and…

Its SKILL.md is about 3.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 49 other files, including reference files and assets (for example `AGENTS.md`, `assets/templates/_template.md` and `metadata.json`).

It sits in Data & Analytics, covering DataFrames and CSV and tabular files. It works with DuckDB and Polars. The repository describes itself as: A collection of AI agent skills following the Agent Skills open format. The licence is MIT.

When your agent uses it

  • Process a large file
  • Code with pd.readcsv of multi-GB files
  • Requests.get(...).content on big bodies
  • BytesIO on unbounded inputs

Example prompts

  • “process a large file”
  • “stream this”
  • “out-of-core”
  • “/io-bound-data-processing”

Workflow steps

9 steps, taken from the step headings in SKILL.md.

  1. Memory Discipline (CRITICAL)
  2. I/O Access Patterns (CRITICAL)
  3. Data Format & Encoding (HIGH)
  4. Chunking & Batching (HIGH)
  5. Spill-to-Disk & External Memory (HIGH)
  6. Pipelining & Backpressure (MEDIUM-HIGH)
  7. Compression & Serialization (MEDIUM)
  8. Concurrency for I/O-Bound Workloads (MEDIUM)
  9. Observability & Throughput Tuning (LOW-MEDIUM)

What it can do on your machine

Read from SKILL.md and the folder at commit cf93c57. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • man7.org
    • github.com
    • arrow.apache.org
    • docs.pola.rs
    • duckdb.org
    • pandas.pydata.org
    • brendangregg.com
    • dataintensive.net

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Io Bound Data Processing loads about 3.3k tokens when it runs, and up to ~37k if it reads all its reference files. Until then it costs about 240 tokens; SKILL.md has 933 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~240
When it runs · the whole SKILL.md, loaded when a task matches
~3.3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~37k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from pproenca/dot-skills at commit cf93c57, republished under its MIT licence (© pproenca). 933 words, ~3,344 tokens.

Download SKILL.mdSave it as .claude/skills/io-bound-data-processing/SKILL.md (or your agent's skills folder). This skill also uses 46 other files; get the full folder from GitHub.
name
io-bound-data-processing
description
Processing, transforming, or moving datasets that may exceed RAM on a single low-compute box — covers memory discipline (streaming, generators, dtype shrinkage), I/O access patterns (sequential vs random, mmap, async), data formats (Parquet vs CSV vs JSON, predicate pushdown), chunking & batching, spill-to-disk (external merge sort, DuckDB/Polars), pipelining (bounded queues, backpressure, checkpointing), codec selection (zstd/lz4/gzip), concurrency for I/O-bound workloads (asyncio, threads, prefetch), and observability (iowait vs CPU%, rows/sec, py-spy/strace). Trigger on "process a large file", "stream this", "out-of-core", "OOM kill", "this is slow", or code with `pd.read_csv` of multi-GB files, `requests.get(...).content` on big bodies, `BytesIO` on unbounded inputs, per-row INSERTs, sequential `requests.get` loops, falling `tqdm` rates — even if I/O or memory isn't mentioned. Complement to computer-science-algorithms.

Community I/O-bound data processing on constrained resources Best Practices

A reference for engineers processing datasets larger than RAM on a single low-compute box. Organized by execution-lifecycle impact: rules near the top of the table govern whether the job runs at all; rules near the bottom shave the last 10 %. Optimize from the top of the waterfall.

Scope: the patterns that show up in real ETL / data-engineering / batch work on a laptop, a 2-vCPU container, or a Raspberry Pi-class node — streaming, formats, chunking, spill, backpressure, codecs, and the concurrency model that actually matches an I/O-bound bottleneck. Out of scope (covered elsewhere): the algorithmic primitives themselves (see computer-science-algorithms), distributed compute beyond a single box (use Spark/Dask), and database-engine internals (see official docs).

Distilled from Apache Arrow / Parquet docs, Polars User Guide, DuckDB docs, pandas — Scaling to large datasets, Linux man pages (mmap(2), sendfile(2), posix_fadvise(2)), Brendan Gregg's USE method and Systems Performance, Kleppmann's Designing Data-Intensive Applications, and the zstd / lz4 reference benchmarks.

When to Apply

Reach for these rules when:

  • A job OOM-kills, swaps, or runs much slower than expected on a small box
  • Input is larger than RAM and you need to scan, filter, aggregate, sort, or join it
  • A pipeline has unbounded buffers between stages, or memory grows linearly during a "streaming" job
  • You see one-row-per-RTT writes (INSERT per row, requests.get per URL, f.read(32) per record)
  • You're picking a format/codec/serializer and the choice matters at scale
  • A top shows low CPU and high iowait, or you don't know which it is
  • "It's slow but I don't know why" — start at the obs- category

Rule Categories By Priority

#CategoryPrefixImpactWhy it cascades
1Memory Disciplinemem-CRITICALSlurping into RAM defeats every downstream technique on a constrained box
2I/O Access Patternsio-CRITICALDisk/net are 10³–10⁶× slower than RAM; access pattern dominates wall-clock
3Data Format & Encodingfmt-HIGHFormat fixes the lower bound on I/O volume + decode cost before any logic runs
4Chunking & Batchingbatch-HIGHGranularity controls peak memory and amortizes per-item overhead
5Spill-to-Disk & External Memoryspill-HIGHWhen data > RAM, the choice is "spill cleanly" or "OOM"
6Pipelining & Backpressurepipe-MEDIUM-HIGHUnbounded buffers between fast producers and slow sinks = OOM
7Compression & Serializationcodec-MEDIUMTrades CPU for I/O; right codec saves orders of magnitude
8Concurrency for I/O-Bound Workloadsconc-MEDIUMAsync / threads / processes are different tools; wrong model wastes CPU
9Observability & Throughput Tuningobs-LOW-MEDIUMCan't tune what you don't measure; iowait ≠ CPU-bound

Quick Reference

1. Memory Discipline (CRITICAL)
2. I/O Access Patterns (CRITICAL)
3. Data Format & Encoding (HIGH)
Show full SKILL.md (379 more words)Show less
4. Chunking & Batching (HIGH)
5. Spill-to-Disk & External Memory (HIGH)
6. Pipelining & Backpressure (MEDIUM-HIGH)
7. Compression & Serialization (MEDIUM)
8. Concurrency for I/O-Bound Workloads (MEDIUM)
9. Observability & Throughput Tuning (LOW-MEDIUM)

How to Use

Start with the question that matches the problem:

Code examples are in Python (most readable across audiences). The reasoning generalizes — equivalent libraries in other ecosystems (Arrow C++/Rust/Java, Polars Rust, DuckDB everywhere, libuv-style async) follow the same patterns.

Reference Files

FileDescription
references/_sections.mdCategory definitions and ordering
assets/templates/_template.mdTemplate for adding new rules
metadata.jsonVersion and reference information
AGENTS.mdAuto-built TOC navigation
  • computer-science-algorithms — Algorithmic primitives this skill builds on (external merge sort, sketches, hash partitioning, sampling)

© pproenca, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 46 other files (references, assets) in skills/.experimental/io-bound-data-processing of pproenca/dot-skills.

  • SKILL.md
  • AGENTS.md
  • assets/templates/_template.md
  • metadata.json
  • references/_sections.md
  • references/batch-coalesce-writes-with-buffered-output.md
  • references/batch-keyset-pagination-over-offset.md
  • references/batch-pick-chunk-size-by-memory-budget.md
  • references/batch-process-with-stable-iterators.md
  • references/batch-use-vectorized-apis-not-row-loops.md
  • references/codec-dictionary-encoding-for-repetitive-strings.md
  • references/codec-prefer-binary-protocols-over-json-for-rpc.md
  • references/codec-train-a-zstd-dictionary-for-many-small-payloads.md
  • references/codec-zstd-or-lz4-as-defaults-not-gzip.md
  • references/conc-asyncio-for-many-network-streams-not-for-cpu.md
  • references/conc-overlap-compute-with-prefetch.md
  • references/conc-thread-pools-for-blocking-io-libraries.md
  • references/conc-tune-parallelism-to-the-bottleneck.md
  • … and 29 more

Open the folder on GitHubat commit cf93c57

Compare with similar skills

Io Bound Data Processing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Io Bound Data Processing compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Io Bound Data Processing this skillpproenca/dot-skills214—~3.3kAutomated safety check: PassMIT
Hybrid-Engine Data Analysiscode-yeongyu/oh-my-openagent70k—~1.4kAutomated safety check: PassCustom licence
Education Data Querybrycewang-stanford/Auto-Empirical-Research-Skills4.5k—~2.8kAutomated safety check: PassCustom licence
Duckdb Experttheneoai/awesome-skills183—~4.2kAutomated safety check: PassMIT
Duckdb EnaAAaqwq/AGI-Super-Team1051 repos~1.6kAutomated safety check: PassMIT
Excel and CSV Data Analysisbytedance/deer-flow83k4 repos~2.2kAutomated safety check: PassMIT

Similar skills

  • Hybrid-Engine Data Analysis

    code-yeongyu/oh-my-openagent

    Analyzes CSV, Parquet and JSON data with DuckDB, Polars, numpy and matplotlib, preferring a persistent kernel over repeated one-shot processes.

    70k GitHub stars~1.4k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Education Data Query

    brycewang-stanford/Auto-Empirical-Research-Skills

    Downloads education datasets from configured mirror sources (parquet/CSV) with local Polars filtering.

    4.5k GitHub stars~2.8k tokensUpdated 2 days ago
    Data & AnalyticsAuto-check passed
  • Duckdb Expert

    theneoai/awesome-skills

    DuckDB expert for embedded OLAP analytics, Parquet/CSV querying, and high-performance analytical SQL on local data.

    183 GitHub stars~4.2k tokensUpdated 4 mo ago
    Data & AnalyticsAuto-check passed
  • Duckdb En

    aAAaqwq/AGI-Super-Team

    DuckDB CLI specialist for SQL analysis, data processing and file conversion.

    105 GitHub starsUsed in 1 repo~1.6k tokens
    DatabasesAuto-check passed
  • Excel and CSV Data Analysis

    bytedance/deer-flow

    Analyzes uploaded Excel and CSV files with SQL through DuckDB, producing schema inspections, statistical summaries and exports to CSV, JSON or Markdown.

    83k GitHub starsUsed in 4 repos~2.2k tokens
    Data & AnalyticsAuto-check passed
  • CSV Data Summarizer

    coffeefuelbump/csv-data-summarizer-claude-skill

    Analyzes CSV files, generates summary stats, and plots quick visualizations using Python and pandas.

    468 GitHub starsUsed in 2 repos~1.4k tokens
    Data & AnalyticsAuto-check passed

More from pproenca/dot-skills

All 182 skills in this repo
  • Audio Voice Recovery

    pproenca/dot-skills

    Audio forensics and voice recovery guidelines for CSI-level audio analysis.

    214 GitHub stars~3.3k tokensUpdated 1 mo ago
    Auto-check passed
  • Codemod React Pipeline

    pproenca/dot-skills

    Guided, scripted pipeline for running JSX/TSX/React codemods safely across large legacy codebases.

    214 GitHub stars~1.6k tokensUpdated 1 mo ago
    Auto-check passed
  • Dev Rfc

    pproenca/dot-skills

    Create well-structured RFCs and technical proposals for software projects.

    214 GitHub stars~3.8k tokensUpdated 1 mo ago
    Auto-check passed
  • Dx Harness

    pproenca/dot-skills

    Developer-experience friction auditing and fixing — slow onboarding, repeated manual setup steps, missing bootstrap/reset/seed scripts, undiscoverable conventions.

    214 GitHub stars~1.5k tokensUpdated 1 mo ago
    Auto-check passed
  • Language Spec Author

    pproenca/dot-skills

    Turn a rough idea for a language into a complete, implementable specification — a DSL, query, config/data, template, or protocol language — by interviewing the author dimension by dimension until…

    214 GitHub stars~2.4k tokensUpdated 1 mo ago
    Auto-check passed
  • Python Pep Author

    pproenca/dot-skills

    Drafting Python Enhancement Proposals (PEPs) — proposing a Python language feature, a standard library change, an interoperability standard, or an informational/process document for the Python…

    214 GitHub stars~2.1k tokensUpdated 1 mo ago
    Auto-check passed

Works with

Questions about Io Bound Data Processing

What does Io Bound Data Processing do?

Processing, transforming, or moving datasets that may exceed RAM on a single low-compute box — covers memory discipline (streaming, generators, dtype shrinkage), I/O access patterns (sequential vs…. Io Bound Data Processing is an agent skill from pproenca/dot-skills.

When should I use Io Bound Data Processing?

Io Bound Data Processing fits situations like: process a large file; code with pd.readcsv of multi-GB files; requests.get(...).content on big bodies; bytesIO on unbounded inputs.

How do I install Io Bound Data Processing in Claude Code?

Run `npx skills add pproenca/dot-skills --skill io-bound-data-processing -a claude-code`. Or copy the skill folder (skills/.experimental/io-bound-data-processing in pproenca/dot-skills) into .claude/skills/io-bound-data-processing in your project. Claude Code loads it when a task matches its description.

How do I install Io Bound Data Processing in Codex?

Run `npx skills add pproenca/dot-skills --skill io-bound-data-processing -a codex`. Or copy the skill folder (skills/.experimental/io-bound-data-processing in pproenca/dot-skills) into .agents/skills/io-bound-data-processing in your project. Codex loads it when a task matches its description.

Can I use Io Bound Data Processing in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add pproenca/dot-skills --skill io-bound-data-processing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/io-bound-data-processing, .gemini/skills/io-bound-data-processing, .github/skills/io-bound-data-processing and .opencode/skills/io-bound-data-processing in your project.

What does Io Bound Data Processing need to run?

SKILL.md names no scripts, command-line tools or credentials: Io Bound Data Processing is instructions for the agent only.

Does Io Bound Data Processing access the network?

SKILL.md names 8 domains. As links in the text: man7.org, github.com, arrow.apache.org, docs.pola.rs, duckdb.org, pandas.pydata.org, brendangregg.com and dataintensive.net. This is read from the text; nothing was executed.

Is Io Bound Data Processing safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Io Bound Data Processing use?

Io Bound Data Processing is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Io Bound Data Processing use?

About 3.3k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 34k tokens, read only when the agent opens those files.

What are the alternatives to Io Bound Data Processing?

Skills that share tags, products or a category with Io Bound Data Processing: Hybrid-Engine Data Analysis (code-yeongyu/oh-my-openagent, 70k stars), Education Data Query (brycewang-stanford/Auto-Empirical-Research-Skills, 4.5k stars), Duckdb Expert (theneoai/awesome-skills, 183 stars) and Duckdb En (aAAaqwq/AGI-Super-Team, 105 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Io Bound Data Processing?

pproenca (a GitHub user) maintains it in pproenca/dot-skills, which has 214 GitHub stars. The repository holds 182 skills in this directory. The repository was last updated on August 15, 2026.

Source: pproenca/dot-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.