Agent skill

Heavy File Ingestion

by NateBJones-Projects in NateBJones-Projects/OB1

Converts large PDF, DOCX, PPTX, XLSX and CSV files into markdown or CSV plus an index before the agent reads them, so tokens go to the compressed copy.

Custom licenceAuto-check passedDocuments & Office

Install Heavy File Ingestion

skills CLI
$ npx skills add NateBJones-Projects/OB1 --skill heavy-file-ingestion -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NateBJones-Projects/OB1 heavy-file-ingestion --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NateBJones-Projects/OB1.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/heavy-file-ingestion .claude/skills/heavy-file-ingestion && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
heavy-file-ingestion
GitHub stars
4.7k
Token cost
~995 tokens
SKILL.md length
479 words
Files
7 (incl. scripts, references)
Skills in repo
24
Repo updated
First seen
Licence
Custom licence

At a glance

Converts large PDF, DOCX, PPTX, XLSX and CSV files into markdown or CSV plus an index before the agent reads them, so tokens go to the compressed copy.

  • Works in 4 steps: Convert before reading. Do not dump raw… → Index before reasoning. Read the… → Match the converter to the file type. → …
  • Asked to read or summarize a large PDF, slide deck or spreadsheet
  • SKILL.md covers Problem, Trigger Conditions, Core Policy and Process, plus 2 more sections
  • Runs Python scripts from its folder; calls python and uv

What it does

Instead of reading a heavyweight file raw, the agent first runs a converter script and reads the generated index.md (or index.json), which says what the file contains, how clean the extraction was and whether escalation is justified. PDFs and documents become a markdown artifact, presentations a markdown slide outline, and spreadsheets one CSV per sheet plus a markdown manifest. Only the outputs the index calls worth reading are opened afterwards.

Escalation follows cost tiers: a deterministic converter with an index first, a cheap model on the extracted artifact only when quality flags show structure was lost, and an expensive model only once the file has been compressed or sampled. The scripts are convert_heavy_file.py and build_client_exports.py, run through uv with libraries such as pdfplumber and python-docx, and the converter can be told to prefer markitdown for PDF and DOCX when it is installed. A reference file covers the open-source stack.

When your agent uses it

  • Asked to read or summarize a large PDF, slide deck or spreadsheet
  • Keeping token use down when a file is bulky or structured
  • Getting a markdown working copy or CSV extraction plus a quick map of a file

Example prompts

  • “Read ./reports/annual-report.pdf and summarize the main findings.”
  • “Look through ./data/orders.xlsx and tell me which sheets matter.”
  • “Summarize the deck in ./pitch/funding-round.pptx without loading the whole file into context.”

Requirements

  • Python with uv to run the converter scripts

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Convert before reading. Do not dump raw heavyweight files into model context if a deterministic converter can create a cheaper artifact.
  2. Index before reasoning. Read the generated index.md or index.json first. It should tell you what is in the file, how clean the extraction…
  3. Match the converter to the file type.
  4. Escalate by cost tier, not instinct.

What it can do on your machine

Read from SKILL.md and the folder at commit 238df6c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python
    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Heavy File Ingestion loads about 995 tokens when it runs, and up to ~1.7k if it reads all its reference files. Until then it costs about 107 tokens; SKILL.md has 479 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~107
When it runs · the whole SKILL.md, loaded when a task matches
~995
With references · SKILL.md plus every file in references/, read only if the agent opens them
~1.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

Its licence (Custom licence) doesn't allow us to republish the file, so here is its outline and opening line. It has 479 words (~995 tokens).

“Agents waste money and context when they read heavyweight files raw. This skill turns bulky documents into cheaper working artifacts first, then tells the main agent how much reasoning power the file actually deserves.”

— opening of SKILL.md by NateBJones-Projects, Custom licence
name
heavy-file-ingestion
author
Nate B. Jones
version
1.0.0

Read the full SKILL.md on GitHub

Files

SKILL.md and 6 other files (scripts, references) in skills/heavy-file-ingestion of NateBJones-Projects/OB1.

  • SKILL.md
  • README.md
  • metadata.json
  • references/open-source-stack.md
  • scripts/build_client_exports.py
  • scripts/convert_heavy_file.py
  • variants

Open the folder on GitHubat commit 238df6c

Compare with similar skills

Heavy File Ingestion next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Heavy File Ingestion compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Heavy File Ingestion this skillNateBJones-Projects/OB14.7k—~995Automated safety check: PassCustom licence
File Format Conversionpipeshub-ai/pipeshub-ai3.8k—~865Automated safety check: PassApache-2.0
Office ArtifactsPrismer-AI/PrismerCloud1.6k—~2.6kAutomated safety check: PassMIT
Document Converterwentorai/Research-Claw858—~1.3kAutomated safety check: PassCustom licence
MarkitdownImCa0/just-laws78114 repos~3.2kAutomated safety check: NotesMIT
MinerU Document Readeropendatalab/MinerU81k—~9.4kAutomated safety check: WarnCustom licence

Similar skills

  • File Format Conversion

    pipeshub-ai/pipeshub-ai

    Picks the right library for converting between CSV, XLSX, JSON, images, DOCX and PDF text, and lists the conversions that are not supported so the agent does not attempt them.

    3.8k GitHub stars~865 tokensUpdated today
    Documents & OfficeAuto-check passed
  • Office Artifacts

    Prismer-AI/PrismerCloud

    Generate real DOCX, PPTX, XLSX, PDF, CSV files using python-docx / python-pptx / openpyxl / reportlab by writing them into the dispatch artifacts dir, then explicitly deliver each one with cloud…

    1.6k GitHub stars~2.6k tokensUpdated today
    Documents & OfficeAuto-check passed
  • Document Converter

    wentorai/Research-Claw

    Convert Office documents (PPTX, DOCX, XLSX, PDF, HTML, CSV, JSON, XML, images) to Markdown using Microsoft MarkItDown.

    858 GitHub stars~1.3k tokensUpdated 1 mo ago
    Documents & OfficeAuto-check passed
  • Markitdown

    ImCa0/just-laws

    Convert files and office documents to Markdown. An agent skill from ImCa0/just-laws.

    781 GitHub starsUsed in 14 repos~3.2k tokens
    Documents & OfficeAuto-check: notes
  • MinerU Document Reader

    opendatalab/MinerU

    Reads, OCRs, searches and cites local documents through the mineru CLI, covering PDF, images, Office files, EPUB, HTML and CSV.

    81k GitHub stars~9.4k tokensUpdated today
    Documents & OfficeAuto-check: warnings
  • Huashu Markdown Publishing Pipeline

    alchaincyf/huashu-md-html

    Converts files and web pages into clean Markdown, then turns Markdown into polished HTML, Word, PDF and EPUB using four templates.

    910 GitHub stars~4.8k tokensUpdated 1 mo ago
    Documents & OfficeAuto-check passed

More from NateBJones-Projects/OB1

All 24 skills in this repo
  • Aiception Skill Extraction

    NateBJones-Projects/OB1

    Pulls reusable knowledge out of work sessions and turns it into new skills, checking existing notes and skills first to avoid duplicates.

    4.7k GitHub stars~2k tokensUpdated 2 days ago
    Auto-check: notes
  • Claude Code Auto-Capture Hook

    NateBJones-Projects/OB1

    Fires a Stop hook that captures a Claude Code session transcript automatically when the session ends without a verbal wrap-up.

    4.7k GitHub stars~1.3k tokensUpdated 2 days ago
    Auto-check: notes
  • Open Brain Local HTTP

    NateBJones-Projects/OB1

    Captures and searches personal notes in a self-hosted Open Brain using plain curl calls over HTTP, for setups where MCP is disabled or blocked.

    4.7k GitHub stars~1.2k tokensUpdated 2 days ago
    Auto-check: notes
  • Panning for Gold

    NateBJones-Projects/OB1

    Processes voice transcripts and brain dumps in three phases, extracting every idea thread, evaluating the strongest and saving a permanent inventory and synthesis.

    4.7k GitHub stars~5k tokensUpdated 2 days ago
    Auto-check passed
  • Agentic Harness Design and Review

    NateBJones-Projects/OB1

    Designs, evaluates and improves the harness around an AI agent: tool permissions, approval gates, state, memory, evals and observability, with phased plans.

    4.7k GitHub stars~1.8k tokensUpdated 2 days ago
    Auto-check passed
  • Weekly Signal Diff

    NateBJones-Projects/OB1

    Turns a noisy week of AI or software market news into a short list of structural changes, weighted by what the user already tracks in Open Brain memory.

    4.7k GitHub stars~1.7k tokensUpdated 2 days ago
    Auto-check passed

Questions about Heavy File Ingestion

What does Heavy File Ingestion do?

Converts large PDF, DOCX, PPTX, XLSX and CSV files into markdown or CSV plus an index before the agent reads them, so tokens go to the compressed copy. json), which says what the file contains, how clean the extraction was and whether escalation is justified. PDFs and documents become a markdown artifact, presentations a markdown slide outline, and spreadsheets one CSV per sheet plus a markdown manifest.

When should I use Heavy File Ingestion?

Heavy File Ingestion fits situations like: asked to read or summarize a large PDF, slide deck or spreadsheet; keeping token use down when a file is bulky or structured; getting a markdown working copy or CSV extraction plus a quick map of a file.

How do I install Heavy File Ingestion in Claude Code?

Run `npx skills add NateBJones-Projects/OB1 --skill heavy-file-ingestion -a claude-code`. Or copy the skill folder (skills/heavy-file-ingestion in NateBJones-Projects/OB1) into .claude/skills/heavy-file-ingestion in your project. Claude Code loads it when a task matches its description.

How do I install Heavy File Ingestion in Codex?

Run `npx skills add NateBJones-Projects/OB1 --skill heavy-file-ingestion -a codex`. Or copy the skill folder (skills/heavy-file-ingestion in NateBJones-Projects/OB1) into .agents/skills/heavy-file-ingestion in your project. Codex loads it when a task matches its description.

Can I use Heavy File Ingestion in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NateBJones-Projects/OB1 --skill heavy-file-ingestion -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/heavy-file-ingestion, .gemini/skills/heavy-file-ingestion, .github/skills/heavy-file-ingestion and .opencode/skills/heavy-file-ingestion in your project.

What does Heavy File Ingestion need to run?

Going by SKILL.md and its folder, Heavy File Ingestion needs Python for the scripts in its folder and the command-line tools its instructions call (python and uv). Our summary lists: Python with uv to run the converter scripts.

Does Heavy File Ingestion access the network?

SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Heavy File Ingestion safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Heavy File Ingestion use?

Heavy File Ingestion has a licence file (the repository's licence) that doesn't match a standard licence. Read it on GitHub before reusing the skill.

How many tokens does Heavy File Ingestion use?

About 995 tokens (SKILL.md is roughly 4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 723 tokens, read only when the agent opens those files.

What are the alternatives to Heavy File Ingestion?

Skills that share tags, products or a category with Heavy File Ingestion: File Format Conversion (pipeshub-ai/pipeshub-ai, 3.8k stars), Office Artifacts (Prismer-AI/PrismerCloud, 1.6k stars), Document Converter (wentorai/Research-Claw, 858 stars) and Markitdown (ImCa0/just-laws, 781 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Heavy File Ingestion?

NateBJones-Projects (a GitHub organization) maintains it in NateBJones-Projects/OB1, which has 4,710 GitHub stars. The repository holds 24 skills in this directory. The repository was last updated on October 9, 2026.

Source: NateBJones-Projects/OB1 on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.