A skill your agent uses when working with structured data files (CSV, JSON, YAML, TOML, Parquet) — querying, transforming, filtering, aggregating, or converting between formats

MITAuto-check passedDocuments & Office

Install Data Processing

skills CLI
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill data-processing -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jeremylongshore/tons-of-skills-marketplace data-processing --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/productivity/cli-power-skills/skills/data-processing .claude/skills/data-processing && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
data-processing
GitHub stars
2.8k
Token cost
~1.3k tokens
SKILL.md length
440 words
Files
1
Skills in repo
3,342
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when working with structured data files (CSV, JSON, YAML, TOML, Parquet) — querying, transforming, filtering, aggregating, or converting between formats

  • Working with structured data files (CSV
  • SKILL.md covers When to Use, Tools, Patterns and Pipelines, plus 2 more sections
  • Calls duckdb, jq and yq
  • Parquet) — querying

What it does

Data Processing is an agent skill from jeremylongshore/tons-of-skills-marketplace. Use when working with structured data files (CSV, JSON, YAML, TOML, Parquet) — querying, transforming, filtering, aggregating, or converting between formats

Its SKILL.md is about 1.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Documents & Office, covering CSV and tabular files, DataFrames and SQL. It works with SQL and DuckDB. The repository describes itself as: Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com. The licence is MIT.

When your agent uses it

  • Working with structured data files (CSV
  • Parquet) — querying
  • Converting between formats

Example prompts

  • “/data-processing”

Requirements

  • Python 3
  • Pre-approved tools (allowed-tools): Bash(jq*), Bash(yq*), Bash(gron*), Bash(mlr*), Bash(xsv*), Bash(duckdb*), Read, Glob, Grep

What it can do on your machine

Read from SKILL.md and the folder at commit 23ea8d4. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Bash(jq*)
    • Bash(yq*)
    • Bash(gron*)
    • Bash(mlr*)
    • Bash(xsv*)
    • Bash(duckdb*)
    • Read
    • Glob
    • Grep

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • duckdb
    • jq
    • yq

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Data Processing loads about 1.3k tokens when it runs. Until then it costs about 43 tokens; SKILL.md has 440 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~43
When it runs · the whole SKILL.md, loaded when a task matches
~1.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jeremylongshore/tons-of-skills-marketplace at commit 23ea8d4, republished under its MIT licence (© jeremylongshore). 440 words, ~1,289 tokens.

Download SKILL.mdSave it as .claude/skills/data-processing/SKILL.md (or your agent's skills folder).
name
data-processing
description
Use when working with structured data files (CSV, JSON, YAML, TOML, Parquet) — querying, transforming, filtering, aggregating, or converting between formats
allowed-tools
Bash(jq*), Bash(yq*), Bash(gron*), Bash(mlr*), Bash(xsv*), Bash(duckdb*), Read, Glob, Grep
version
1.0.0
author
ykotik
license
MIT

Data Processing

When to Use

  • Querying, filtering, or transforming JSON files
  • Reading or converting YAML, TOML, or XML config files
  • Analyzing, aggregating, or joining CSV/TSV/Parquet files
  • Running SQL queries against local data files without a database server
  • Converting between data formats (JSON to YAML, CSV to JSON, etc.)
  • Exploring deeply nested JSON structures

Tools

ToolPurposeStructured output
jqQuery and transform JSONNative JSON
yqQuery and transform YAML, TOML, XML-o json for JSON output
gronFlatten JSON into greppable path.to.key = value lines--ungron to reverse back to JSON
miller (mlr)Transform CSV/JSON/TSV records with awk-like verbs--json for JSON output
xsvFast CSV slicing, searching, joining, statisticsCSV native (pipe to xsv table for display)
DuckDBSQL queries on CSV, JSON, Parquet files-json flag for JSON output

Patterns

JSON: Filter array elements by field value
bash
jq '.items[] | select(.status == "active")' data.json
JSON: Extract specific fields from array
bash
jq '[.users[] | {name: .name, email: .email}]' data.json
JSON: Count items grouped by field
bash
jq '[.events[] | .type] | group_by(.) | map({type: .[0], count: length})' data.json
YAML: Read a nested value
bash
yq '.spec.containers[0].image' deployment.yaml
YAML: Convert entire file to JSON
bash
yq -o json config.yaml
TOML: Read a value
bash
yq -p toml '.database.host' config.toml
JSON: Explore unknown structure by grepping paths
bash
gron data.json | grep -i "error"
CSV: Column statistics (min, max, mean, stddev)
bash
xsv stats data.csv | xsv table
CSV: Search rows matching a pattern in a column
bash
xsv search -s status "active" data.csv
CSV: Select specific columns
bash
xsv select name,email,created_at users.csv
CSV: Sort by column
bash
xsv sort -s revenue -R sales.csv
SQL: Query a CSV file
bash
duckdb -c "SELECT department, COUNT(*) as cnt, AVG(salary) as avg_sal FROM 'employees.csv' GROUP BY department ORDER BY avg_sal DESC"
SQL: Query a JSON file
bash
duckdb -c "SELECT * FROM read_json_auto('events.json') WHERE type = 'error' LIMIT 20"
SQL: Query Parquet files
bash
duckdb -c "SELECT * FROM 'data/*.parquet' WHERE created_at > '2026-01-01'"
CSV/JSON: Transform records with miller
bash
mlr --csv filter '$revenue > 1000' then sort-by -nr revenue sales.csv
Format conversion: CSV to JSON
bash
mlr --icsv --ojson cat data.csv

Pipelines

YAML config → JSON → SQL query
bash
yq -o json config.yaml | duckdb -c "SELECT key, value FROM read_json_auto('/dev/stdin') WHERE env = 'production'"

Each stage: yq converts YAML to JSON, DuckDB runs SQL on the JSON stream.

Grep nested JSON paths → reconstruct matching subset
bash
gron large.json | grep "\.errors\[" | gron --ungron

Each stage: gron flattens JSON to paths, grep filters, ungron reconstructs valid JSON from matches.

Show full SKILL.md (177 more words)Show less
CSV filter → aggregate with SQL
bash
xsv search -s region "EU" sales.csv | duckdb -c "SELECT product, SUM(revenue) as total FROM read_csv_auto('/dev/stdin') GROUP BY product ORDER BY total DESC"

Each stage: xsv filters rows by region, DuckDB aggregates the filtered stream.

Join two CSV files with SQL
bash
duckdb -c "SELECT u.name, u.email, o.total, o.date FROM 'users.csv' u JOIN 'orders.csv' o ON u.id = o.user_id ORDER BY o.date DESC"
Multi-format pipeline: JSON → CSV → stats
bash
jq -r '.records[] | [.name, .score] | @csv' data.json | xsv stats

Each stage: jq extracts fields to CSV format, xsv computes statistics.

Prefer Over

  • Prefer DuckDB over Python/pandas for ad-hoc SQL queries on files — single command, no script, handles large files
  • Prefer jq over Python json module for one-off JSON transforms — single pipeline vs. multi-line script
  • Prefer xsv over awk/cut for CSV operations — correct CSV parsing, handles quoted fields and escapes
  • Prefer miller over awk for format-aware record transformations — understands CSV/JSON headers natively
  • Prefer yq over custom parsers for config file reads — handles YAML, TOML, XML with consistent syntax

Do NOT Use When

  • Data is already in a running database — query the database directly
  • File is under 10 lines — just use the Read tool and process in-context
  • Task requires complex multi-step logic with conditionals — write a Python script instead
  • JSON is simple enough to read by eye — use the Read tool, don't over-engineer
  • Working with binary or non-structured data formats

© jeremylongshore, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in plugins/productivity/cli-power-skills/skills/data-processing of jeremylongshore/tons-of-skills-marketplace.

Open the folder on GitHubat commit 23ea8d4

Compare with similar skills

Data Processing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Data Processing compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Data Processing this skilljeremylongshore/tons-of-skills-marketplace2.8k—~1.3kAutomated safety check: PassMIT
Duckdb EnaAAaqwq/AGI-Super-Team1051 repos~1.6kAutomated safety check: PassMIT
Matlab Use Duckdbmatlab/matlab-agentic-toolkit1.1k—~3.9kAutomated safety check: PassCustom licence
Data AnalysisHezaoHezao/poirot250—~997Automated safety check: PassMIT
Excel and CSV Data Analysisbytedance/deer-flow83k4 repos~2.2kAutomated safety check: PassMIT
Chdb SQLvemetric/vemetric3941 repos~1.2kAutomated safety check: PassApache-2.0

Similar skills

  • Duckdb En

    aAAaqwq/AGI-Super-Team

    DuckDB CLI specialist for SQL analysis, data processing and file conversion.

    105 GitHub starsUsed in 1 repo~1.6k tokens
    DatabasesAuto-check passed
  • Matlab Use Duckdb

    matlab/matlab-agentic-toolkit

    Use DuckDB from MATLAB via Database Toolbox (R2026a+) as a non-math operations engine on large tabular files (CSV/Parquet/JSON) and as a zero-config embedded database.

    1.1k GitHub stars~3.9k tokensUpdated today
    DatabasesAuto-check passed
  • Data Analysis

    HezaoHezao/poirot

    Analyze Excel/CSV files with DuckDB SQL via bash. An agent skill from HezaoHezao/poirot.

    250 GitHub stars~997 tokensUpdated 2 mo ago
    Documents & OfficeAuto-check passed
  • Excel and CSV Data Analysis

    bytedance/deer-flow

    Analyzes uploaded Excel and CSV files with SQL through DuckDB, producing schema inspections, statistical summaries and exports to CSV, JSON or Markdown.

    83k GitHub starsUsed in 4 repos~2.2k tokens
    Data & AnalyticsAuto-check passed
  • Chdb SQL

    vemetric/vemetric

    A skill your agent uses when the user wants to run SQL — especially analytical SQL — on local files (parquet/csv/json), URLs, S3 paths, or remote databases (Postgres, MySQL, MongoDB, ClickHouse…

    394 GitHub starsUsed in 1 repo~1.2k tokens
    DatabasesAuto-check passed
  • Ops Telemetry Query

    boundless-xyz/boundless

    Internal — for Boundless team members only. An agent skill from boundless-xyz/boundless.

    193 GitHub stars~3.8k tokensUpdated 1 mo ago
    DatabasesAuto-check passed

More from jeremylongshore/tons-of-skills-marketplace

All 3,342 skills in this repo
  • Performing Security Code Review

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to conduct a security-focused code review using the security-agent plugin.

    2.8k GitHub starsUsed in 2 repos~1.3k tokens
    Auto-check: notes
  • Analyzing Text With NLP

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to perform natural language processing and text analysis using the nlp-text-analyzer plugin.

    2.8k GitHub starsUsed in 1 repo~819 tokens
    Auto-check passed
  • Building Neural Networks

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill allows AI assistant to construct and configure neural network architectures using the neural-network-builder plugin.

    2.8k GitHub starsUsed in 1 repo~1k tokens
    Auto-check passed
  • Detecting Data Anomalies

    jeremylongshore/tons-of-skills-marketplace

    Process identify anomalies and outliers in datasets using machine learning algorithms.

    2.8k GitHub starsUsed in 1 repo~1.4k tokens
    Auto-check passed
  • Explaining Machine Learning Models

    jeremylongshore/tons-of-skills-marketplace

    Build this skill enables AI assistant to provide interpretability and explainability for machine learning models.

    2.8k GitHub starsUsed in 1 repo~1k tokens
    Auto-check passed
  • Optimizing Prompts

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill optimizes prompts for large language models (llms) to reduce token usage, lower costs, and improve performance.

    2.8k GitHub starsUsed in 1 repo~1k tokens
    Auto-check passed

Works with

Questions about Data Processing

What does Data Processing do?

A skill your agent uses when working with structured data files (CSV, JSON, YAML, TOML, Parquet) — querying, transforming, filtering, aggregating, or converting between formats. Data Processing is an agent skill from jeremylongshore/tons-of-skills-marketplace.

When should I use Data Processing?

Data Processing fits situations like: working with structured data files (CSV; parquet) — querying; converting between formats.

How do I install Data Processing in Claude Code?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill data-processing -a claude-code`. Or copy the skill folder (plugins/productivity/cli-power-skills/skills/data-processing in jeremylongshore/tons-of-skills-marketplace) into .claude/skills/data-processing in your project. Claude Code loads it when a task matches its description.

How do I install Data Processing in Codex?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill data-processing -a codex`. Or copy the skill folder (plugins/productivity/cli-power-skills/skills/data-processing in jeremylongshore/tons-of-skills-marketplace) into .agents/skills/data-processing in your project. Codex loads it when a task matches its description.

Can I use Data Processing in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill data-processing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/data-processing, .gemini/skills/data-processing, .github/skills/data-processing and .opencode/skills/data-processing in your project.

What does Data Processing need to run?

Going by SKILL.md and its folder, Data Processing needs the command-line tools its instructions call (duckdb, jq and yq). Our summary lists: Python 3. Its frontmatter pre-approves these tools: Bash(jq*), Bash(yq*), Bash(gron*), Bash(mlr*), Bash(xsv*), Bash(duckdb*), Read, Glob, Grep.

Does Data Processing access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Data Processing safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Data Processing use?

Data Processing is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Data Processing use?

About 1.3k tokens (SKILL.md is roughly 5.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Data Processing?

Skills that share tags, products or a category with Data Processing: Duckdb En (aAAaqwq/AGI-Super-Team, 105 stars), Matlab Use Duckdb (matlab/matlab-agentic-toolkit, 1.1k stars), Data Analysis (HezaoHezao/poirot, 250 stars) and Excel and CSV Data Analysis (bytedance/deer-flow, 83k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Data Processing?

jeremylongshore (a GitHub user) maintains it in jeremylongshore/tons-of-skills-marketplace, which has 2,821 GitHub stars. The repository holds 3,342 skills in this directory. The repository was last updated on October 8, 2026.

Source: jeremylongshore/tons-of-skills-marketplace on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.