Agent skill

CSV Pipeline

by aAAaqwq in aAAaqwq/AGI-Super-Team

Process, transform, analyze, and report on CSV and JSON data files.

MITAuto-check passedDocuments & Office

Install CSV Pipeline

skills CLI
$ npx skills add aAAaqwq/AGI-Super-Team --skill csv-pipeline -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install aAAaqwq/AGI-Super-Team csv-pipeline --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/aAAaqwq/AGI-Super-Team.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/csv-pipeline .claude/skills/csv-pipeline && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
csv-pipeline
GitHub stars
105
Used in
1 other repo
Token cost
~3.1k tokens
SKILL.md length
199 words
Files
2
Skills in repo
167
Repo updated
First seen
Licence
MIT

At a glance

Process, transform, analyze, and report on CSV and JSON data files.

  • The user needs to filter rows
  • SKILL.md covers When to Use, Quick Operations with Standard…, Python Operations (for complex… and Format Conversion, plus 4 more sections
  • Calls sqlite3
  • Compute aggregates

What it does

CSV Pipeline is an agent skill from aAAaqwq/AGI-Super-Team. Process, transform, analyze, and report on CSV and JSON data files. Use when the user needs to filter rows, join datasets, compute aggregates, convert formats, deduplicate, or generate summary reports from tabular data. Works with any CSV, TSV, or JSON Lines file.

Its SKILL.md is about 3.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `_meta.json`).

It sits in Documents & Office, covering CSV and tabular files. It works with Python. The repository describes itself as: An installable, cross-framework AI organization: C-suite agents, expert subagents, curated skills, independent review, and one-command setup across 18 AI client/runtime adapters. The licence is MIT.

When your agent uses it

  • The user needs to filter rows
  • Compute aggregates
  • Convert formats
  • Generate summary reports from tabular data

Example prompts

  • “/csv-pipeline”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 331ecd3. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • sqlite3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

CSV Pipeline loads about 3.1k tokens when it runs. Until then it costs about 69 tokens; SKILL.md has 199 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~69
When it runs · the whole SKILL.md, loaded when a task matches
~3.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from aAAaqwq/AGI-Super-Team at commit 331ecd3, republished under its MIT licence (© aAAaqwq). 199 words, ~3,144 tokens.

Download SKILL.mdSave it as .claude/skills/csv-pipeline/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
csv-pipeline
description
Process, transform, analyze, and report on CSV and JSON data files. Use when the user needs to filter rows, join datasets, compute aggregates, convert formats, deduplicate, or generate summary reports from tabular data. Works with any CSV, TSV, or JSON Lines file.

CSV Data Pipeline

Process tabular data (CSV, TSV, JSON, JSON Lines) using standard command-line tools and Python. No external dependencies required beyond Python 3.

When to Use

  • User provides a CSV/TSV/JSON file and asks to analyze, transform, or report on it
  • Joining, filtering, grouping, or aggregating tabular data
  • Converting between formats (CSV to JSON, JSON to CSV, etc.)
  • Deduplicating, sorting, or cleaning messy data
  • Generating summary statistics or reports
  • ETL workflows: extract from one format, transform, load into another

Quick Operations with Standard Tools

Inspect
bash
# Preview first rows
head -5 data.csv

# Count rows (excluding header)
tail -n +2 data.csv | wc -l

# Show column headers
head -1 data.csv

# Count unique values in a column (column 3)
tail -n +2 data.csv | cut -d',' -f3 | sort -u | wc -l
Filter with awk
bash
# Filter rows where column 3 > 100
awk -F',' 'NR==1 || $3 > 100' data.csv > filtered.csv

# Filter rows matching a pattern in column 2
awk -F',' 'NR==1 || $2 ~ /pattern/' data.csv > matched.csv

# Sum column 4
awk -F',' 'NR>1 {sum += $4} END {print sum}' data.csv
Sort and Deduplicate
bash
# Sort by column 2 (numeric)
head -1 data.csv > sorted.csv && tail -n +2 data.csv | sort -t',' -k2 -n >> sorted.csv

# Deduplicate by all columns
head -1 data.csv > deduped.csv && tail -n +2 data.csv | sort -u >> deduped.csv

# Deduplicate by specific column (keep first occurrence)
awk -F',' '!seen[$2]++' data.csv > deduped.csv

Python Operations (for complex transforms)

Read and Inspect
python
import csv, json, sys
from collections import Counter

def read_csv(path, delimiter=','):
    """Read CSV/TSV into list of dicts."""
    with open(path, newline='', encoding='utf-8') as f:
        return list(csv.DictReader(f, delimiter=delimiter))

def write_csv(rows, path, delimiter=','):
    """Write list of dicts to CSV."""
    if not rows:
        return
    with open(path, 'w', newline='', encoding='utf-8') as f:
        writer = csv.DictWriter(f, fieldnames=rows[0].keys(), delimiter=delimiter)
        writer.writeheader()
        writer.writerows(rows)

# Quick stats
data = read_csv('data.csv')
print(f"Rows: {len(data)}")
print(f"Columns: {list(data[0].keys())}")
for col in data[0]:
    non_empty = sum(1 for r in data if r[col].strip())
    print(f"  {col}: {non_empty}/{len(data)} non-empty")
Filter and Transform
python
# Filter rows
filtered = [r for r in data if float(r['amount']) > 100]

# Add computed column
for r in data:
    r['total'] = str(float(r['price']) * int(r['quantity']))

# Rename columns
renamed = [{('new_name' if k == 'old_name' else k): v for k, v in r.items()} for r in data]

# Type conversion
for r in data:
    r['amount'] = float(r['amount'])
    r['date'] = r['date'].strip()
Group and Aggregate
python
from collections import defaultdict

def group_by(rows, key):
    """Group rows by a column value."""
    groups = defaultdict(list)
    for r in rows:
        groups[r[key]].append(r)
    return dict(groups)

def aggregate(rows, group_col, agg_col, func='sum'):
    """Aggregate a column by groups."""
    groups = group_by(rows, group_col)
    results = []
    for name, group in sorted(groups.items()):
        values = [float(r[agg_col]) for r in group if r[agg_col].strip()]
        if func == 'sum':
            agg = sum(values)
        elif func == 'avg':
            agg = sum(values) / len(values) if values else 0
        elif func == 'count':
            agg = len(values)
        elif func == 'min':
            agg = min(values) if values else 0
        elif func == 'max':
            agg = max(values) if values else 0
        results.append({group_col: name, f'{func}_{agg_col}': str(agg), 'count': str(len(group))})
    return results

# Example: sum revenue by category
summary = aggregate(data, 'category', 'revenue', 'sum')
write_csv(summary, 'summary.csv')
Join Datasets
python
def inner_join(left, right, on):
    """Inner join two datasets on a key column."""
    right_index = {}
    for r in right:
        key = r[on]
        if key not in right_index:
            right_index[key] = []
        right_index[key].append(r)

    results = []
    for lr in left:
        key = lr[on]
        if key in right_index:
            for rr in right_index[key]:
                merged = {**lr}
                for k, v in rr.items():
                    if k != on:
                        merged[k] = v
                results.append(merged)
    return results

def left_join(left, right, on):
    """Left join: keep all left rows, fill missing right with empty."""
    right_index = {}
    right_cols = set()
    for r in right:
        key = r[on]
        right_cols.update(r.keys())
        if key not in right_index:
            right_index[key] = []
        right_index[key].append(r)
    right_cols.discard(on)

    results = []
    for lr in left:
        key = lr[on]
        if key in right_index:
            for rr in right_index[key]:
                merged = {**lr}
                for k, v in rr.items():
                    if k != on:
                        merged[k] = v
                results.append(merged)
        else:
            merged = {**lr}
            for col in right_cols:
                merged[col] = ''
            results.append(merged)
    return results

# Example
orders = read_csv('orders.csv')
customers = read_csv('customers.csv')
joined = left_join(orders, customers, on='customer_id')
write_csv(joined, 'orders_with_customers.csv')
Deduplicate
python
def deduplicate(rows, key_cols=None):
    """Remove duplicate rows. If key_cols specified, dedupe by those columns only."""
    seen = set()
    unique = []
    for r in rows:
        if key_cols:
            key = tuple(r[c] for c in key_cols)
        else:
            key = tuple(sorted(r.items()))
        if key not in seen:
            seen.add(key)
            unique.append(r)
    return unique

# Deduplicate by email column
clean = deduplicate(data, key_cols=['email'])

Format Conversion

CSV to JSON
python
import json, csv

with open('data.csv', newline='', encoding='utf-8') as f:
    rows = list(csv.DictReader(f))

# Array of objects
with open('data.json', 'w') as f:
    json.dump(rows, f, indent=2)

# JSON Lines (one object per line, streamable)
with open('data.jsonl', 'w') as f:
    for row in rows:
        f.write(json.dumps(row) + '\n')
JSON to CSV
python
import json, csv

with open('data.json') as f:
    rows = json.load(f)

with open('data.csv', 'w', newline='', encoding='utf-8') as f:
    writer = csv.DictWriter(f, fieldnames=rows[0].keys())
    writer.writeheader()
    writer.writerows(rows)
JSON Lines to CSV
python
import json, csv

rows = []
with open('data.jsonl') as f:
    for line in f:
        if line.strip():
            rows.append(json.loads(line))

with open('data.csv', 'w', newline='', encoding='utf-8') as f:
    all_keys = set()
    for r in rows:
        all_keys.update(r.keys())
    writer = csv.DictWriter(f, fieldnames=sorted(all_keys))
    writer.writeheader()
    writer.writerows(rows)
TSV to CSV
bash
tr '\t' ',' < data.tsv > data.csv

Data Cleaning Patterns

Fix common CSV issues
python
def clean_csv(rows):
    """Clean common CSV data quality issues."""
    cleaned = []
    for r in rows:
        clean_row = {}
        for k, v in r.items():
            # Strip whitespace from keys and values
            k = k.strip()
            v = v.strip() if isinstance(v, str) else v
            # Normalize empty values
            if v in ('', 'N/A', 'n/a', 'NA', 'null', 'NULL', 'None', '-'):
                v = ''
            # Normalize boolean values
            if v.lower() in ('true', 'yes', '1', 'y'):
                v = 'true'
            elif v.lower() in ('false', 'no', '0', 'n'):
                v = 'false'
            clean_row[k] = v
        cleaned.append(clean_row)
    return cleaned
Validate data types
python
def validate_rows(rows, schema):
    """
    Validate rows against a schema.
    schema: dict of column_name -> 'int'|'float'|'date'|'email'|'str'
    Returns (valid_rows, error_rows)
    """
    import re
    valid, errors = [], []
    for i, r in enumerate(rows):
        errs = []
        for col, dtype in schema.items():
            val = r.get(col, '').strip()
            if not val:
                continue
            if dtype == 'int':
                try:
                    int(val)
                except ValueError:
                    errs.append(f"{col}: '{val}' not int")
            elif dtype == 'float':
                try:
                    float(val)
                except ValueError:
                    errs.append(f"{col}: '{val}' not float")
            elif dtype == 'email':
                if not re.match(r'^[^@]+@[^@]+\.[^@]+$', val):
                    errs.append(f"{col}: '{val}' not email")
            elif dtype == 'date':
                if not re.match(r'^\d{4}-\d{2}-\d{2}', val):
                    errs.append(f"{col}: '{val}' not YYYY-MM-DD")
        if errs:
            errors.append({'row': i + 2, 'errors': errs, 'data': r})
        else:
            valid.append(r)
    return valid, errors

# Usage
valid, bad = validate_rows(data, {'amount': 'float', 'email': 'email', 'date': 'date'})
print(f"Valid: {len(valid)}, Errors: {len(bad)}")
for e in bad[:5]:
    print(f"  Row {e['row']}: {e['errors']}")

Generating Reports

Summary report as Markdown
python
def generate_report(data, title, group_col, value_col):
    """Generate a Markdown summary report."""
    lines = [f"# {title}", f"", f"**Total rows**: {len(data)}", ""]

    # Group summary
    groups = group_by(data, group_col)
    lines.append(f"## By {group_col}")
    lines.append("")
    lines.append(f"| {group_col} | Count | Sum | Avg | Min | Max |")
    lines.append("|---|---|---|---|---|---|")

    for name in sorted(groups):
        vals = [float(r[value_col]) for r in groups[name] if r[value_col].strip()]
        if vals:
            lines.append(f"| {name} | {len(vals)} | {sum(vals):.2f} | {sum(vals)/len(vals):.2f} | {min(vals):.2f} | {max(vals):.2f} |")

    lines.append("")
    lines.append(f"*Generated from {len(data)} rows*")
    return '\n'.join(lines)

report = generate_report(data, "Sales Summary", "category", "revenue")
with open('report.md', 'w') as f:
    f.write(report)

Large File Handling

For files too large to load into memory at once:

python
def stream_process(input_path, output_path, transform_fn, delimiter=','):
    """Process a CSV row-by-row without loading entire file."""
    with open(input_path, newline='', encoding='utf-8') as fin, \
         open(output_path, 'w', newline='', encoding='utf-8') as fout:
        reader = csv.DictReader(fin, delimiter=delimiter)
        writer = None
        for row in reader:
            result = transform_fn(row)
            if result is None:
                continue  # Skip row
            if writer is None:
                writer = csv.DictWriter(fout, fieldnames=result.keys(), delimiter=delimiter)
                writer.writeheader()
            writer.writerow(result)

# Example: filter and transform in streaming fashion
def process_row(row):
    if float(row.get('amount', 0) or 0) < 10:
        return None  # Skip small amounts
    row['amount_usd'] = str(float(row['amount']) * 1.0)  # Add computed field
    return row

stream_process('big_file.csv', 'output.csv', process_row)

Tips

  • Always check encoding: file -i data.csv or open with encoding='utf-8-sig' for BOM files
  • For Excel exports with commas in values, the CSV module handles quoting automatically
  • Use json.dumps(ensure_ascii=False) for international characters
  • Pipe-delimited files: use delimiter='|' in csv.reader/writer
  • For very large aggregations, consider sqlite3 which Python includes:
    bash
    sqlite3 :memory: ".mode csv" ".import data.csv t" "SELECT category, SUM(amount) FROM t GROUP BY category;"

© aAAaqwq, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in skills/csv-pipeline of aAAaqwq/AGI-Super-Team.

  • SKILL.md
  • _meta.json

Open the folder on GitHubat commit 331ecd3

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in aAAaqwq/AGI-Super-Team, which our catalogue first saw on October 7, 2026.

Compare with similar skills

CSV Pipeline next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

CSV Pipeline compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
CSV Pipeline this skillaAAaqwq/AGI-Super-Team1051 repos~3.1kAutomated safety check: PassMIT
Instrument Data To Allotropeaws-samples/amazon-bedrock-agents-healthcare-lifesciences2742 repos~2.7kAutomated safety check: PassApache-2.0
Sap Rpt1secondsky/sap-skills462—~1.8kAutomated safety check: NotesGPL-3.0
Replicate Paperbrycewang-stanford/Auto-Empirical-Research-Skills4.5k—~1.9kAutomated safety check: NotesCustom licence
Regression Insightrongxinzy/RongxinAI154—~839Automated safety check: PassMIT
Meta Forest Binary Plotaipoch/medical-research-skills2k—~2.1kAutomated safety check: PassMIT

Similar skills

  • Instrument Data To Allotrope

    aws-samples/amazon-bedrock-agents-healthcare-lifesciences

    Official

    Convert laboratory instrument output files (PDF, CSV, Excel, TXT) to Allotrope Simple Model (ASM) JSON format or flattened 2D CSV.

    274 GitHub starsUsed in 2 repos~2.7k tokens
    Documents & OfficeAuto-check passed
  • Sap Rpt1

    secondsky/sap-skills

    SAP-RPT-1-OSS local tabular prediction workflows for FI/CO prototype datasets.

    462 GitHub stars~1.8k tokensUpdated 4 days ago
    Documents & OfficeAuto-check: notes
  • Replicate Paper

    brycewang-stanford/Auto-Empirical-Research-Skills

    Run a full 6-phase autonomous replication of a biomedical/epidemiology paper against UK Biobank or similar cohort data, producing Python and R scripts plus a validated replication report.

    4.5k GitHub stars~1.9k tokensUpdated 4 days ago
    Documents & OfficeAuto-check: notes
  • Regression Insight

    rongxinzy/RongxinAI

    对 CSV/Excel 数据执行线性回归(OLS)或逻辑回归(Logistic),一键输出完整统计结果(包含回归系数、R²、p值、VIF等)和中文通俗解读。当用户提及回归分析、拟合模型、查看系数显著性、R方、p值、共线性(VIF),或使用关键词如 回归、regression、OLS、logit、拟合、显著性 时触发。

    154 GitHub stars~839 tokensUpdated today
    Documents & OfficeAuto-check passed
  • Meta Forest Binary Plot

    aipoch/medical-research-skills

    Generate meta-analysis forest plots for binary classification data.

    2k GitHub stars~2.1k tokensUpdated 22 days ago
    Documents & OfficeAuto-check passed
  • Discogs Sync

    LeoYeAI/openclaw-master-skills

    Add and remove albums from a Discogs wantlist or collection by artist and album name, master ID, or release ID.

    2.2k GitHub stars~3.9k tokensUpdated 2 mo ago
    Documents & OfficeAuto-check passed

More from aAAaqwq/AGI-Super-Team

All 167 skills in this repo
  • Content Creator

    aAAaqwq/AGI-Super-Team

    Create SEO-optimized marketing content with consistent brand voice.

    105 GitHub starsUsed in 3 repos~1.9k tokens
    Auto-check passed
  • Financial Calculator

    aAAaqwq/AGI-Super-Team

    Advanced financial calculator with future value tables, present value, discount calculations, markup pricing, and compound interest.

    105 GitHub starsUsed in 1 repo~1.5k tokens
    Auto-check passed
  • Bankr Signals

    aAAaqwq/AGI-Super-Team

    Transaction-verified trading signals on Base blockchain. An agent skill from aAAaqwq/AGI-Super-Team.

    105 GitHub starsUsed in 2 repos~3.3k tokens
    Auto-check passed
  • Erc 8004

    aAAaqwq/AGI-Super-Team

    Register AI agents on Ethereum mainnet using ERC-8004 (Trustless Agents).

    105 GitHub starsUsed in 2 repos~1.2k tokens
    Auto-check passed
  • Frontend Design Ultimate

    aAAaqwq/AGI-Super-Team

    Create distinctive, production-grade static sites with React, Tailwind CSS, and shadcn/ui — no mockups needed.

    105 GitHub starsUsed in 2 repos~2.7k tokens
    Auto-check passed
  • Zsxq Smart Publish

    aAAaqwq/AGI-Super-Team

    Publish and manage content on 知识星球 (zsxq.com). An agent skill from aAAaqwq/AGI-Super-Team.

    105 GitHub stars~1.5k tokensUpdated today
    Auto-check passed

Works with

Questions about CSV Pipeline

What does CSV Pipeline do?

Process, transform, analyze, and report on CSV and JSON data files. CSV Pipeline is an agent skill from aAAaqwq/AGI-Super-Team. Process, transform, analyze, and report on CSV and JSON data files.

When should I use CSV Pipeline?

CSV Pipeline fits situations like: the user needs to filter rows; compute aggregates; convert formats; generate summary reports from tabular data.

How do I install CSV Pipeline in Claude Code?

Run `npx skills add aAAaqwq/AGI-Super-Team --skill csv-pipeline -a claude-code`. Or copy the skill folder (skills/csv-pipeline in aAAaqwq/AGI-Super-Team) into .claude/skills/csv-pipeline in your project. Claude Code loads it when a task matches its description.

How do I install CSV Pipeline in Codex?

Run `npx skills add aAAaqwq/AGI-Super-Team --skill csv-pipeline -a codex`. Or copy the skill folder (skills/csv-pipeline in aAAaqwq/AGI-Super-Team) into .agents/skills/csv-pipeline in your project. Codex loads it when a task matches its description.

Can I use CSV Pipeline in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add aAAaqwq/AGI-Super-Team --skill csv-pipeline -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/csv-pipeline, .gemini/skills/csv-pipeline, .github/skills/csv-pipeline and .opencode/skills/csv-pipeline in your project.

What does CSV Pipeline need to run?

Going by SKILL.md and its folder, CSV Pipeline needs the command-line tools its instructions call (sqlite3). Our summary lists: Python 3.

Does CSV Pipeline access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is CSV Pipeline safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does CSV Pipeline use?

CSV Pipeline is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does CSV Pipeline use?

About 3.1k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to CSV Pipeline?

Skills that share tags, products or a category with CSV Pipeline: Instrument Data To Allotrope (aws-samples/amazon-bedrock-agents-healthcare-lifesciences, 274 stars), Sap Rpt1 (secondsky/sap-skills, 462 stars), Replicate Paper (brycewang-stanford/Auto-Empirical-Research-Skills, 4.5k stars) and Regression Insight (rongxinzy/RongxinAI, 154 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains CSV Pipeline?

aAAaqwq (a GitHub user) maintains it in aAAaqwq/AGI-Super-Team, which has 105 GitHub stars. The repository holds 167 skills in this directory. The repository was last updated on October 8, 2026.

Source: aAAaqwq/AGI-Super-Team on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.