Hybrid-Engine Data Analysis
code-yeongyu/oh-my-openagent
Analyzes CSV, Parquet and JSON data with DuckDB, Polars, numpy and matplotlib, preferring a persistent kernel over repeated one-shot processes.
Processing, transforming, or moving datasets that may exceed RAM on a single low-compute box — covers memory discipline (streaming, generators, dtype shrinkage), I/O access patterns (sequential vs…
$ npx skills add pproenca/dot-skills --skill io-bound-data-processing -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install pproenca/dot-skills io-bound-data-processing --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/pproenca/dot-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/.experimental/io-bound-data-processing .claude/skills/io-bound-data-processing && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "io-bound-data-processing" agent skill from https://github.com/pproenca/dot-skills/tree/master/skills/.experimental/io-bound-data-processing into .claude/skills/io-bound-data-processing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "io-bound-data-processing", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/pproenca/dot-skills/tree/master/skills/.experimental/io-bound-data-processingType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add pproenca/dot-skills --skill io-bound-data-processing -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install pproenca/dot-skills io-bound-data-processing --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/pproenca/dot-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/.experimental/io-bound-data-processing .agents/skills/io-bound-data-processing && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "io-bound-data-processing" agent skill from https://github.com/pproenca/dot-skills/tree/master/skills/.experimental/io-bound-data-processing into .agents/skills/io-bound-data-processing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "io-bound-data-processing", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add pproenca/dot-skills --skill io-bound-data-processing -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install pproenca/dot-skills io-bound-data-processing --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/pproenca/dot-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/.experimental/io-bound-data-processing .cursor/skills/io-bound-data-processing && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "io-bound-data-processing" agent skill from https://github.com/pproenca/dot-skills/tree/master/skills/.experimental/io-bound-data-processing into .cursor/skills/io-bound-data-processing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "io-bound-data-processing", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/pproenca/dot-skills.git --path skills/.experimental/io-bound-data-processing--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add pproenca/dot-skills --skill io-bound-data-processing -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install pproenca/dot-skills io-bound-data-processing --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/pproenca/dot-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/.experimental/io-bound-data-processing .gemini/skills/io-bound-data-processing && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "io-bound-data-processing" agent skill from https://github.com/pproenca/dot-skills/tree/master/skills/.experimental/io-bound-data-processing into .gemini/skills/io-bound-data-processing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "io-bound-data-processing", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install pproenca/dot-skills io-bound-data-processingInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add pproenca/dot-skills --skill io-bound-data-processing -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/pproenca/dot-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/.experimental/io-bound-data-processing .github/skills/io-bound-data-processing && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "io-bound-data-processing" agent skill from https://github.com/pproenca/dot-skills/tree/master/skills/.experimental/io-bound-data-processing into .github/skills/io-bound-data-processing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "io-bound-data-processing", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add pproenca/dot-skills --skill io-bound-data-processing -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install pproenca/dot-skills io-bound-data-processing --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/pproenca/dot-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/.experimental/io-bound-data-processing .opencode/skills/io-bound-data-processing && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "io-bound-data-processing" agent skill from https://github.com/pproenca/dot-skills/tree/master/skills/.experimental/io-bound-data-processing into .opencode/skills/io-bound-data-processing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "io-bound-data-processing", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
io-bound-data-processingProcessing, transforming, or moving datasets that may exceed RAM on a single low-compute box — covers memory discipline (streaming, generators, dtype shrinkage), I/O access patterns (sequential vs…
Io Bound Data Processing is an agent skill from pproenca/dot-skills. Processing, transforming, or moving datasets that may exceed RAM on a single low-compute box — covers memory discipline (streaming, generators, dtype shrinkage), I/O access patterns (sequential vs random, mmap, async), data formats (Parquet vs CSV vs JSON, predicate pushdown), chunking & batching, spill-to-disk (external merge sort, DuckDB/Polars), pipelining (bounded queues, backpressure, checkpointing), codec selection (zstd/lz4/gzip), concurrency for I/O-bound workloads (asyncio, threads, prefetch), and…
Its SKILL.md is about 3.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 49 other files, including reference files and assets (for example `AGENTS.md`, `assets/templates/_template.md` and `metadata.json`).
It sits in Data & Analytics, covering DataFrames and CSV and tabular files. It works with DuckDB and Polars. The repository describes itself as: A collection of AI agent skills following the Agent Skills open format. The licence is MIT.
9 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit cf93c57. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md.
From the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
man7.orggithub.comarrow.apache.orgdocs.pola.rsduckdb.orgpandas.pydata.orgbrendangregg.comdataintensive.netFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Io Bound Data Processing loads about 3.3k tokens when it runs, and up to ~37k if it reads all its reference files. Until then it costs about 240 tokens; SKILL.md has 933 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from pproenca/dot-skills at commit cf93c57, republished under its MIT licence (© pproenca). 933 words, ~3,344 tokens.
.claude/skills/io-bound-data-processing/SKILL.md (or your agent's skills folder). This skill also uses 46 other files; get the full folder from GitHub.A reference for engineers processing datasets larger than RAM on a single low-compute box. Organized by execution-lifecycle impact: rules near the top of the table govern whether the job runs at all; rules near the bottom shave the last 10 %. Optimize from the top of the waterfall.
Scope: the patterns that show up in real ETL / data-engineering / batch work on a laptop, a 2-vCPU container, or a Raspberry Pi-class node — streaming, formats, chunking, spill, backpressure, codecs, and the concurrency model that actually matches an I/O-bound bottleneck. Out of scope (covered elsewhere): the algorithmic primitives themselves (see computer-science-algorithms), distributed compute beyond a single box (use Spark/Dask), and database-engine internals (see official docs).
Distilled from Apache Arrow / Parquet docs, Polars User Guide, DuckDB docs, pandas — Scaling to large datasets, Linux man pages (mmap(2), sendfile(2), posix_fadvise(2)), Brendan Gregg's USE method and Systems Performance, Kleppmann's Designing Data-Intensive Applications, and the zstd / lz4 reference benchmarks.
Reach for these rules when:
INSERT per row, requests.get per URL, f.read(32) per record)top shows low CPU and high iowait, or you don't know which it is| # | Category | Prefix | Impact | Why it cascades |
|---|---|---|---|---|
| 1 | Memory Discipline | mem- | CRITICAL | Slurping into RAM defeats every downstream technique on a constrained box |
| 2 | I/O Access Patterns | io- | CRITICAL | Disk/net are 10³–10⁶× slower than RAM; access pattern dominates wall-clock |
| 3 | Data Format & Encoding | fmt- | HIGH | Format fixes the lower bound on I/O volume + decode cost before any logic runs |
| 4 | Chunking & Batching | batch- | HIGH | Granularity controls peak memory and amortizes per-item overhead |
| 5 | Spill-to-Disk & External Memory | spill- | HIGH | When data > RAM, the choice is "spill cleanly" or "OOM" |
| 6 | Pipelining & Backpressure | pipe- | MEDIUM-HIGH | Unbounded buffers between fast producers and slow sinks = OOM |
| 7 | Compression & Serialization | codec- | MEDIUM | Trades CPU for I/O; right codec saves orders of magnitude |
| 8 | Concurrency for I/O-Bound Workloads | conc- | MEDIUM | Async / threads / processes are different tools; wrong model wastes CPU |
| 9 | Observability & Throughput Tuning | obs- | LOW-MEDIUM | Can't tune what you don't measure; iowait ≠ CPU-bound |
mem-stream-dont-slurp — Iterate sources chunk-by-chunk; peak RAM = chunk size, not file sizemem-prefer-generators-over-lists-for-pipelines — Generators flow; lists materializemem-shrink-dtypes-before-loading — Narrow ints, categoricals; 2-8× memory reduction at load timemem-use-views-not-copies — Slicing without copying; NumPy/Arrow zero-copy semanticsmem-bound-the-working-set — Chunk size = budget ÷ row-size × amplification, not a round numbermem-release-references-explicitly — Drop intermediates so peak ≠ N × chunkio-prefer-sequential-over-random — Sort offsets, advise the kernel, let readahead helpio-buffer-explicitly-for-small-records — BufferedReader collapses 1000× syscallsio-stream-http-bodies-with-iter-content — stream=True + iter_content instead of .contentio-mmap-for-random-or-shared-large-files — Zero-copy + on-demand paging for random accessio-async-for-many-concurrent-streams — One thread, thousands of awaitsio-batch-and-pipeline-network-roundtrips — COPY, pipelines, multi-key endpoints, HTTP/2io-zero-copy-when-moving-bytes-as-is — sendfile, copy_file_range, shutil.copyfilefmt-columnar-for-analytical-scans — Parquet/Arrow for filter+project workloadsfmt-line-delimited-for-streaming-row-ingest — NDJSON over JSON-array for streamingfmt-push-predicates-into-the-reader — Row-group statistics skip whole chunksfmt-prefer-schema-on-write-when-possible — Typed columns beat schema-on-read every timefmt-avoid-deeply-nested-json-for-hot-paths — Flat schema, or binary on hot pipesbatch-pick-chunk-size-by-memory-budget — Compute from budget, not a constantbatch-use-vectorized-apis-not-row-loops — NumPy / Polars / Arrow kernels, not iterrowsbatch-process-with-stable-iterators — iter_batches, chunksize=, collect(streaming=True)batch-coalesce-writes-with-buffered-output — COPY / executemany, sized write buffersbatch-keyset-pagination-over-offset — WHERE id > $last_id, never OFFSET N on deep cursorsspill-external-merge-sort-when-data-exceeds-ram — Out-of-core sort; delegate to sort / DuckDB when possiblespill-partition-by-hash-for-out-of-core-groupby-join — Hash-partition both sides; process per partitionspill-use-temp-files-not-process-memory — SpooledTemporaryFile, never unbounded BytesIOspill-use-engines-that-spill-automatically — DuckDB / Polars / Dask manage spill for youpipe-use-bounded-queues-for-producer-consumer — Bounded Queue is the backpressure mechanismpipe-apply-backpressure-from-slow-stages — Slow sink throttles the fast sourcepipe-prefer-pull-iteration-over-push-callbacks — Pull backpressures naturally; push needs policypipe-checkpoint-progress-for-resumability — Atomic checkpoint after each batch; redo boundedcodec-zstd-or-lz4-as-defaults-not-gzip — Pick codec by access pattern; gzip is legacycodec-dictionary-encoding-for-repetitive-strings — 5-50× on low-cardinality columnscodec-prefer-binary-protocols-over-json-for-rpc — Protobuf / Arrow IPC / MsgPack on hot wirescodec-train-a-zstd-dictionary-for-many-small-payloads — zstd --train for sub-1 KB messagesconc-asyncio-for-many-network-streams-not-for-cpu — Asyncio multiplexes waits; useless for computeconc-thread-pools-for-blocking-io-libraries — GIL releases during blocking I/O; threads workconc-overlap-compute-with-prefetch — One batch ahead via background thread / coroutineconc-tune-parallelism-to-the-bottleneck — Match worker count to the binding resourceobs-measure-iowait-not-just-cpu — iostat -x, vmstat, psutil — find the real bottleneckobs-instrument-throughput-rows-per-second — Rates normalize across runs; catch regressionsobs-profile-with-py-spy-or-strace-for-syscall-storms — Attach, don't guessStart with the question that matches the problem:
mem- and pipe- (likely an unbounded buffer or per-iteration accumulation)io- (likely random access, unbuffered I/O, or wrong format)fmt-columnar-for-analytical-scans and io-buffer-explicitly-for-small-recordsbatch-pick-chunk-size-by-memory-budgetspill- (and prefer an engine that does it for you: spill-use-engines-that-spill-automatically)io-batch-and-pipeline-network-roundtrips, io-async-for-many-concurrent-streamsconc-tune-parallelism-to-the-bottleneck, obs-measure-iowait-not-just-cpuobs- first; profile before changing anythingCode examples are in Python (most readable across audiences). The reasoning generalizes — equivalent libraries in other ecosystems (Arrow C++/Rust/Java, Polars Rust, DuckDB everywhere, libuv-style async) follow the same patterns.
| File | Description |
|---|---|
| references/_sections.md | Category definitions and ordering |
| assets/templates/_template.md | Template for adding new rules |
| metadata.json | Version and reference information |
| AGENTS.md | Auto-built TOC navigation |
computer-science-algorithms — Algorithmic primitives this skill builds on (external merge sort, sketches, hash partitioning, sampling)© pproenca, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 46 other files (references, assets) in skills/.experimental/io-bound-data-processing of pproenca/dot-skills.
Open the folder on GitHubat commit cf93c57
Io Bound Data Processing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Io Bound Data Processing this skillpproenca/dot-skills | 214 | — | ~3.3k | Automated safety check: Pass | MIT | |
| Hybrid-Engine Data Analysiscode-yeongyu/oh-my-openagent | 70k | — | ~1.4k | Automated safety check: Pass | Custom licence | |
| Education Data Querybrycewang-stanford/Auto-Empirical-Research-Skills | 4.5k | — | ~2.8k | Automated safety check: Pass | Custom licence | |
| Duckdb Experttheneoai/awesome-skills | 183 | — | ~4.2k | Automated safety check: Pass | MIT | |
| Duckdb EnaAAaqwq/AGI-Super-Team | 105 | 1 repos | ~1.6k | Automated safety check: Pass | MIT | |
| Excel and CSV Data Analysisbytedance/deer-flow | 83k | 4 repos | ~2.2k | Automated safety check: Pass | MIT |
code-yeongyu/oh-my-openagent
Analyzes CSV, Parquet and JSON data with DuckDB, Polars, numpy and matplotlib, preferring a persistent kernel over repeated one-shot processes.
brycewang-stanford/Auto-Empirical-Research-Skills
Downloads education datasets from configured mirror sources (parquet/CSV) with local Polars filtering.
theneoai/awesome-skills
DuckDB expert for embedded OLAP analytics, Parquet/CSV querying, and high-performance analytical SQL on local data.
aAAaqwq/AGI-Super-Team
DuckDB CLI specialist for SQL analysis, data processing and file conversion.
bytedance/deer-flow
Analyzes uploaded Excel and CSV files with SQL through DuckDB, producing schema inspections, statistical summaries and exports to CSV, JSON or Markdown.
coffeefuelbump/csv-data-summarizer-claude-skill
Analyzes CSV files, generates summary stats, and plots quick visualizations using Python and pandas.
pproenca/dot-skills
Audio forensics and voice recovery guidelines for CSI-level audio analysis.
pproenca/dot-skills
Guided, scripted pipeline for running JSX/TSX/React codemods safely across large legacy codebases.
pproenca/dot-skills
Create well-structured RFCs and technical proposals for software projects.
pproenca/dot-skills
Developer-experience friction auditing and fixing — slow onboarding, repeated manual setup steps, missing bootstrap/reset/seed scripts, undiscoverable conventions.
pproenca/dot-skills
Turn a rough idea for a language into a complete, implementable specification — a DSL, query, config/data, template, or protocol language — by interviewing the author dimension by dimension until…
pproenca/dot-skills
Drafting Python Enhancement Proposals (PEPs) — proposing a Python language feature, a standard library change, an interoperability standard, or an informational/process document for the Python…
Categories
Processing, transforming, or moving datasets that may exceed RAM on a single low-compute box — covers memory discipline (streaming, generators, dtype shrinkage), I/O access patterns (sequential vs…. Io Bound Data Processing is an agent skill from pproenca/dot-skills.
Io Bound Data Processing fits situations like: process a large file; code with pd.readcsv of multi-GB files; requests.get(...).content on big bodies; bytesIO on unbounded inputs.
Run `npx skills add pproenca/dot-skills --skill io-bound-data-processing -a claude-code`. Or copy the skill folder (skills/.experimental/io-bound-data-processing in pproenca/dot-skills) into .claude/skills/io-bound-data-processing in your project. Claude Code loads it when a task matches its description.
Run `npx skills add pproenca/dot-skills --skill io-bound-data-processing -a codex`. Or copy the skill folder (skills/.experimental/io-bound-data-processing in pproenca/dot-skills) into .agents/skills/io-bound-data-processing in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add pproenca/dot-skills --skill io-bound-data-processing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/io-bound-data-processing, .gemini/skills/io-bound-data-processing, .github/skills/io-bound-data-processing and .opencode/skills/io-bound-data-processing in your project.
SKILL.md names no scripts, command-line tools or credentials: Io Bound Data Processing is instructions for the agent only.
SKILL.md names 8 domains. As links in the text: man7.org, github.com, arrow.apache.org, docs.pola.rs, duckdb.org, pandas.pydata.org, brendangregg.com and dataintensive.net. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Io Bound Data Processing is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.3k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 34k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Io Bound Data Processing: Hybrid-Engine Data Analysis (code-yeongyu/oh-my-openagent, 70k stars), Education Data Query (brycewang-stanford/Auto-Empirical-Research-Skills, 4.5k stars), Duckdb Expert (theneoai/awesome-skills, 183 stars) and Duckdb En (aAAaqwq/AGI-Super-Team, 105 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
pproenca (a GitHub user) maintains it in pproenca/dot-skills, which has 214 GitHub stars. The repository holds 182 skills in this directory. The repository was last updated on August 15, 2026.
Source: pproenca/dot-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.