Agent skill

File Format Conversion

by pipeshub-ai in pipeshub-ai/pipeshub-ai

Picks the right library for converting between CSV, XLSX, JSON, images, DOCX and PDF text, and lists the conversions that are not supported so the agent does not attempt them.

Apache-2.0Auto-check passedDocuments & Office

Install File Format Conversion

skills CLI
$ npx skills add pipeshub-ai/pipeshub-ai --skill file-conversion -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install pipeshub-ai/pipeshub-ai file-conversion --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/pipeshub-ai/pipeshub-ai.git skills-src && mkdir -p .claude/skills && cp -r skills-src/backend/python/app/agents/agent_loop/skills/builtin_packs/file-conversion .claude/skills/file-conversion && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
file-conversion
GitHub stars
3.8k
Token cost
~865 tokens
SKILL.md length
345 words
Files
1
Skills in repo
9
Repo updated
First seen
Licence
Apache-2.0

At a glance

Picks the right library for converting between CSV, XLSX, JSON, images, DOCX and PDF text, and lists the conversions that are not supported so the agent does not attempt them.

  • Converting a CSV to XLSX, JSON or HTML
  • SKILL.md covers Supported conversions, NOT supported in this… and Choosing between "extract…
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Pulling text out of a DOCX or tables out of a PDF

What it does

The skill is a dispatch table for file conversions. For csv, xlsx and json it defaults to Node libraries, papaparse for parsing and exceljs for spreadsheets, with a hand-rolled table loop for html output, and it falls back to pandas when a conversion needs real reshaping such as pivots, group-bys or type coercion. Images are converted or resized with sharp, with Pillow as the Python fallback. DOCX text extraction uses python-docx or docx2python, and PDF text and tables use pdfplumber, because no solid Node equivalent exists for those.

A separate list names the conversions that are not supported because the sandbox has no document-rendering engine (LibreOffice): docx to pdf, pptx to pdf, legacy doc to docx, xlsx to pdf, and anything that needs a document layout rendered to an image. The agent must tell you directly instead of faking a result or claiming success. A final section helps decide between extracting data and converting the format.

When your agent uses it

  • Converting a CSV to XLSX, JSON or HTML
  • Pulling text out of a DOCX or tables out of a PDF
  • Checking whether a requested conversion is possible in this environment

Example prompts

  • “Convert sales.csv to an Excel file.”
  • “Extract the tables from report.pdf into a CSV.”
  • “Turn this PowerPoint into a PDF, or tell me if that is not possible here.”

Requirements

  • Node.js with papaparse, exceljs and sharp
  • Python with pandas, Pillow, python-docx and pdfplumber for the fallback paths

What it can do on your machine

Read from SKILL.md and the folder at commit a883af6. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

File Format Conversion loads about 865 tokens when it runs. Until then it costs about 111 tokens; SKILL.md has 345 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~111
When it runs · the whole SKILL.md, loaded when a task matches
~865

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from pipeshub-ai/pipeshub-ai at commit a883af6, republished under its Apache-2.0 licence (© pipeshub-ai). 345 words, ~865 tokens.

Download SKILL.mdSave it as .claude/skills/file-conversion/SKILL.md (or your agent's skills folder).
name
file-conversion
description
Use this skill when the user asks to convert a file from one format to another (e.g. csv to xlsx, xlsx to json, pdf text to csv) and you're not sure which library handles that specific pair, or whether it's supported in this environment at all. Dispatches to the right library per format pair and explicitly lists conversions that are NOT supported here so you don't attempt one that will silently fail or produce a wrong result.

File format conversion

Supported conversions

FromToDefault (Node)Fallback (Python)
csvxlsx / json / htmlpapaparse (parse to array of objects) → exceljs (xlsx) / JSON.stringify (json) / a hand-rolled <table> loop (html — no library needed for this)pandas.read_csv → .to_excel / .to_json / .to_html — reach for this when the conversion needs real reshaping (pivots, group-by, dtype coercion) along the way, not as a first choice
xlsxcsv / json / htmlexceljs (read) → papaparse.unparse (csv) / JSON.stringify (json) / hand-rolled htmlpandas.read_excel → .to_csv / .to_json / .to_html
json (tabular/list-of-records shape)csv / xlsxJSON.parse → papaparse.unparse (csv) / exceljs (xlsx)pandas.read_json → .to_csv / .to_excel
image (png/jpg/webp/etc.)different image format, resize, basic editssharp (.toFormat(...), .resize(...))Pillow (Image.open(...).save(new_path))
docxplain text— (no solid Node equivalent for structured docx parsing)python-docx / docx2python — extract paragraph and table text
pdfplain text / csv (tables)— (no solid Node equivalent for table-aware extraction)pdfplumber — extract_text() / extract_tables() → build a DataFrame if tabular

For plain csv/xlsx/json reshuffling, default to the Node path (an array of plain objects is the common intermediate, playing the same role pandas.DataFrame plays on the Python side) per this environment's Node-first policy — reach for pandas only when the conversion needs real data reshaping along the way. docx/pdf extraction and structure-aware parsing have no solid Node equivalent here, so those stay Python regardless of format on the other end.

NOT supported in this environment

These all require a document-rendering engine (LibreOffice) that is not installed in the sandbox yet. Tell the user directly rather than attempting a workaround (e.g. do not try to fake a "pdf" by screenshotting text, and do not claim success on a conversion you didn't actually perform):

  • docx → pdf
  • pptx → pdf
  • doc (legacy binary Word format) → docx
  • xlsx → pdf
  • Any conversion that requires rendering/rasterizing a document (visual layout to image)

Choosing between "extract data" and "convert format"

If the user's underlying goal is "get the numbers out of this PDF/docx into a spreadsheet" rather than literally "convert this file", extracting the relevant data and writing exactly the columns/rows they need is usually a better outcome than a generic whole-file conversion — ask yourself which one they actually need before picking an approach.

© pipeshub-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in backend/python/app/agents/agent_loop/skills/builtin_packs/file-conversion of pipeshub-ai/pipeshub-ai.

Open the folder on GitHubat commit a883af6

Compare with similar skills

File Format Conversion next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

File Format Conversion compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
File Format Conversion this skillpipeshub-ai/pipeshub-ai3.8k—~865Automated safety check: PassApache-2.0
Office ArtifactsPrismer-AI/PrismerCloud1.6k—~2.6kAutomated safety check: PassMIT
Heavy File IngestionNateBJones-Projects/OB14.7k—~995Automated safety check: PassCustom licence
Streaming Export Safetydoccker/cc-use-exp1.1k—~1.9kAutomated safety check: PassCustom licence
MinerU Document Readeropendatalab/MinerU81k—~9.4kAutomated safety check: WarnCustom licence
ComPDF Image to Document ConverterComPDFKit/compdf-skills109—~891Automated safety check: PassNone

Similar skills

  • Office Artifacts

    Prismer-AI/PrismerCloud

    Generate real DOCX, PPTX, XLSX, PDF, CSV files using python-docx / python-pptx / openpyxl / reportlab by writing them into the dispatch artifacts dir, then explicitly deliver each one with cloud…

    1.6k GitHub stars~2.6k tokensUpdated 9 days ago
    Documents & OfficeAuto-check passed
  • Heavy File Ingestion

    NateBJones-Projects/OB1

    Converts large PDF, DOCX, PPTX, XLSX and CSV files into markdown or CSV plus an index before the agent reads them, so tokens go to the compressed copy.

    4.7k GitHub stars~995 tokensUpdated yesterday
    Documents & OfficeAuto-check passed
  • Streaming Export Safety

    doccker/cc-use-exp

    当代码涉及 Excel/CSV/JSON/PDF 大文件导出、批量序列化、内存里构建大对象时触发。防止 OOM、临时文件残留、同步导出阻塞 HTTP 线程等内存安全陷阱。

    1.1k GitHub stars~1.9k tokensUpdated 1 mo ago
    Documents & OfficeAuto-check passed
  • MinerU Document Reader

    opendatalab/MinerU

    Reads, OCRs, searches and cites local documents through the mineru CLI, covering PDF, images, Office files, EPUB, HTML and CSV.

    81k GitHub stars~9.4k tokensUpdated today
    Documents & OfficeAuto-check: warnings
  • ComPDF Image to Document Converter

    ComPDFKit/compdf-skills

    Plans ComPDF Server API requests that turn screenshots, scans, receipts and forms into Word, Excel, PPT, PDF, HTML, RTF, CSV, TXT or JSON output.

    109 GitHub stars~891 tokensUpdated 1 mo ago
    Documents & OfficeAuto-check passed
  • Multi Source Data Integration Extraction

    Drchronx/ai-agent-research-starter-kit

    Automatically merge scattered Excel and CSV files, normalize column names, and extract structured tables from PDF, HTML, TXT, or Markdown documents.

    137 GitHub stars~671 tokensUpdated 4 mo ago
    Documents & OfficeAuto-check passed

More from pipeshub-ai/pipeshub-ai

All 9 skills in this repo
  • PipesHub Excel Spreadsheet Builder

    pipeshub-ai/pipeshub-ai

    Creates and edits .xlsx workbooks with real Excel formulas rather than hardcoded computed values, defaulting to exceljs in TypeScript with a static formula-safety check.

    3.8k GitHub stars~1.8k tokensUpdated today
    Auto-check passed
  • Office Open XML Utilities

    pipeshub-ai/pipeshub-ai

    Unpacks a .docx or .pptx into pretty-printed XML, lets you make small targeted edits, and repacks it into a file Office will open.

    3.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Picks the right library for generating a new PDF, filling an existing PDF form, or extracting text and tables, defaulting to Node where possible.

    3.8k GitHub stars~2.9k tokensUpdated today
    Auto-check passed
  • PowerPoint Deck Builder

    pipeshub-ai/pipeshub-ai

    Creates new PowerPoint decks with pptxgenjs in TypeScript, reads existing decks with python-pptx, and applies a design-quality checklist so every slide has real visual hierarchy.

    3.8k GitHub stars~1.6k tokensUpdated today
    Auto-check passed
  • Verified Data Analysis with pandas

    pipeshub-ai/pipeshub-ai

    Loads, cleans, aggregates and joins tabular data with pandas under a verification rule: every number reported must be one that the code actually printed.

    3.8k GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Chart Type Selection Guide

    pipeshub-ai/pipeshub-ai

    Picks the right chart type for a data question and applies readability rules like axis labels, colorblind palettes and legend restraint.

    3.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed

Questions about File Format Conversion

What does File Format Conversion do?

Picks the right library for converting between CSV, XLSX, JSON, images, DOCX and PDF text, and lists the conversions that are not supported so the agent does not attempt them. The skill is a dispatch table for file conversions. For csv, xlsx and json it defaults to Node libraries, papaparse for parsing and exceljs for spreadsheets, with a hand-rolled table loop for html output, and it falls back to pandas when a conversion needs real reshaping such as pivots, group-bys or type coercion.

When should I use File Format Conversion?

File Format Conversion fits situations like: converting a CSV to XLSX, JSON or HTML; pulling text out of a DOCX or tables out of a PDF; checking whether a requested conversion is possible in this environment.

How do I install File Format Conversion in Claude Code?

Run `npx skills add pipeshub-ai/pipeshub-ai --skill file-conversion -a claude-code`. Or copy the skill folder (backend/python/app/agents/agent_loop/skills/builtin_packs/file-conversion in pipeshub-ai/pipeshub-ai) into .claude/skills/file-conversion in your project. Claude Code loads it when a task matches its description.

How do I install File Format Conversion in Codex?

Run `npx skills add pipeshub-ai/pipeshub-ai --skill file-conversion -a codex`. Or copy the skill folder (backend/python/app/agents/agent_loop/skills/builtin_packs/file-conversion in pipeshub-ai/pipeshub-ai) into .agents/skills/file-conversion in your project. Codex loads it when a task matches its description.

Can I use File Format Conversion in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add pipeshub-ai/pipeshub-ai --skill file-conversion -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/file-conversion, .gemini/skills/file-conversion, .github/skills/file-conversion and .opencode/skills/file-conversion in your project.

What does File Format Conversion need to run?

SKILL.md names no scripts, command-line tools or credentials: File Format Conversion is instructions for the agent only. Our summary lists: Node.js with papaparse, exceljs and sharp; Python with pandas, Pillow, python-docx and pdfplumber for the fallback paths.

Does File Format Conversion access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is File Format Conversion safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does File Format Conversion use?

File Format Conversion is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does File Format Conversion use?

About 865 tokens (SKILL.md is roughly 3.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to File Format Conversion?

Skills that share tags, products or a category with File Format Conversion: Office Artifacts (Prismer-AI/PrismerCloud, 1.6k stars), Heavy File Ingestion (NateBJones-Projects/OB1, 4.7k stars), Streaming Export Safety (doccker/cc-use-exp, 1.1k stars) and MinerU Document Reader (opendatalab/MinerU, 81k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains File Format Conversion?

pipeshub-ai (a GitHub organization) maintains it in pipeshub-ai/pipeshub-ai, which has 3,821 GitHub stars. The repository holds 9 skills in this directory. The repository was last updated on October 9, 2026.

Source: pipeshub-ai/pipeshub-ai on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.