Agent skill

Heavy File Ingestion Claude Code

by NateBJones-Projects in NateBJones-Projects/OB1

Use in Claude Code when a user asks to read, analyze, summarize, or extract from a heavyweight file such as PDF, DOCX, PPTX, XLSX, CSV, or TSV.

Custom licenceAuto-check passedDocuments & Office

Install Heavy File Ingestion Claude Code

skills CLI
$ npx skills add NateBJones-Projects/OB1 --skill heavy-file-ingestion-claude-code -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NateBJones-Projects/OB1 heavy-file-ingestion-claude-code --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NateBJones-Projects/OB1.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/heavy-file-ingestion/variants/claude-code .claude/skills/heavy-file-ingestion-claude-code && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
heavy-file-ingestion-claude-code
GitHub stars
4.7k
Token cost
~614 tokens
SKILL.md length
263 words
Files
1
Skills in repo
24
Repo updated
First seen
Licence
Custom licence

At a glance

Use in Claude Code when a user asks to read, analyze, summarize, or extract from a heavyweight file such as PDF, DOCX, PPTX, XLSX, CSV, or TSV.

  • Works in 7 steps: Do not read the original heavyweight… → Resolve the bundled converter relative… → Run the converter first. Default command → …
  • Tasks that involve CSV and tabular files
  • SKILL.md covers Problem, Trigger Conditions, Process and Client Rules, plus 1 more section
  • Calls python and uv

What it does

Heavy File Ingestion Claude Code is an agent skill from NateBJones-Projects/OB1. Use in Claude Code when a user asks to read, analyze, summarize, or extract from a heavyweight file such as PDF, DOCX, PPTX, XLSX, CSV, or TSV. Convert the file into markdown or CSV first with the bundled script, generate a lightweight index, and only spend model tokens on the compressed artifact.

Its SKILL.md is about 610 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Documents & Office, covering CSV and tabular files, Excel spreadsheets and Word documents. It works with Microsoft Excel, Microsoft PowerPoint and Microsoft Word. The repository describes itself as: Open Brain — The infrastructure layer for your thinking. One database, one AI gateway, one chat channel — any AI plugs in. No middleware, no SaaS.

When your agent uses it

  • Tasks that involve CSV and tabular files
  • Tasks that involve Excel spreadsheets
  • Tasks that involve Word documents

Example prompts

  • “/heavy-file-ingestion-claude-code”

Requirements

  • Python 3

Workflow steps

7 steps, taken from the first numbered list in SKILL.md.

  1. Do not read the original heavyweight file directly into context if conversion is possible.
  2. Resolve the bundled converter relative to this skill directory: scripts/convert_heavy_file.py
  3. Run the converter first. Default command
  4. If dependencies are missing, prefer
  5. Read the generated index.md before reading any converted artifact.
  6. Use the index to decide the cheapest next step
  7. Only escalate to a stronger model after the file has already been compressed into markdown, CSV, or a short sampled subset.

What it can do on your machine

Read from SKILL.md and the folder at commit 238df6c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python
    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Heavy File Ingestion Claude Code loads about 614 tokens when it runs. Until then it costs about 83 tokens; SKILL.md has 263 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~83
When it runs · the whole SKILL.md, loaded when a task matches
~614

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

Its licence (Custom licence) doesn't allow us to republish the file, so here is its outline and opening line. It has 263 words (~614 tokens).

“Claude Code has the tools to convert files locally, so it should not waste context by reading heavyweight files raw.”

— opening of SKILL.md by NateBJones-Projects, Custom licence
name
heavy-file-ingestion-claude-code
author
Nate B. Jones
version
1.0.0

Read the full SKILL.md on GitHub

Files

Just SKILL.md in skills/heavy-file-ingestion/variants/claude-code of NateBJones-Projects/OB1.

Open the folder on GitHubat commit 238df6c

Compare with similar skills

Heavy File Ingestion Claude Code next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Heavy File Ingestion Claude Code compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Heavy File Ingestion Claude Code this skillNateBJones-Projects/OB14.7k—~614Automated safety check: PassCustom licence
Light File ReadingLight0305/Light-skills640—~4.1kAutomated safety check: PassMIT
Document Converterwentorai/Research-Claw858—~1.3kAutomated safety check: PassCustom licence
Doc SummarizerBlackBeltTechnology/pi-agent-dashboard315—~1.2kAutomated safety check: PassMIT
To MarkdownMathews-Tom/armory329—~2kAutomated safety check: PassMIT
Compdf Conversion CLILeoYeAI/openclaw-master-skills2.2k—~5.9kAutomated safety check: PassProprietary

Similar skills

  • Light File Reading

    Light0305/Light-skills

    Light 多格式文件深度理解常驻技能:强大地读 Word / PDF / PPTX / Excel / CSV / 图片 / 视频 / 代码 / 压缩包,不只提取文字,而是理解结构 / 图表 / 数据 / 格式要求 / 隐含意图,产结构化"理解笔记"五面 (结构逻辑·关键内容·格式约束·视觉风格·可复用)并映射到下游技能动作(这个文件→接下来能做什么)。

    640 GitHub stars~4.1k tokensUpdated 3 mo ago
    Documents & OfficeAuto-check passed
  • Document Converter

    wentorai/Research-Claw

    Convert Office documents (PPTX, DOCX, XLSX, PDF, HTML, CSV, JSON, XML, images) to Markdown using Microsoft MarkItDown.

    858 GitHub stars~1.3k tokensUpdated 1 mo ago
    Documents & OfficeAuto-check passed
  • Doc Summarizer

    BlackBeltTechnology/pi-agent-dashboard

    Summarize documents of any size: extract with the document-converter engine, chunk to fit context, fan out to subagents, then synthesize one unified summary.

    315 GitHub stars~1.2k tokensUpdated today
    Documents & OfficeAuto-check passed
  • To Markdown

    Mathews-Tom/armory

    Convert any file or URL to clean Markdown: PDF, DOCX, XLSX, PPTX, HTML, images (OCR), audio, CSV, YouTube.

    329 GitHub stars~2k tokensUpdated 4 days ago
    Documents & OfficeAuto-check passed
  • Compdf Conversion CLI

    LeoYeAI/openclaw-master-skills

    MUST use for ANY PDF or image format conversion task — converting PDF and images (JPG/JPEG/PNG/BMP/TIFF/TIF/WEBP/JPEG2000) to 10 formats (Word, Excel, PPT, HTML, Image, TXT, JSON, Markdown, RTF…

    2.2k GitHub stars~5.9k tokensUpdated 2 mo ago
    Documents & OfficeAuto-check passed
  • MinerU Document Reader

    opendatalab/MinerU

    Reads, OCRs, searches and cites local documents through the mineru CLI, covering PDF, images, Office files, EPUB, HTML and CSV.

    81k GitHub stars~9.4k tokensUpdated yesterday
    Documents & OfficeAuto-check: warnings

More from NateBJones-Projects/OB1

All 24 skills in this repo
  • Heavy File Ingestion

    NateBJones-Projects/OB1

    Converts large PDF, DOCX, PPTX, XLSX and CSV files into markdown or CSV plus an index before the agent reads them, so tokens go to the compressed copy.

    4.7k GitHub stars~995 tokensUpdated yesterday
    Auto-check passed
  • Aiception Skill Extraction

    NateBJones-Projects/OB1

    Pulls reusable knowledge out of work sessions and turns it into new skills, checking existing notes and skills first to avoid duplicates.

    4.7k GitHub stars~2k tokensUpdated yesterday
    Auto-check: notes
  • Claude Code Auto-Capture Hook

    NateBJones-Projects/OB1

    Fires a Stop hook that captures a Claude Code session transcript automatically when the session ends without a verbal wrap-up.

    4.7k GitHub stars~1.3k tokensUpdated yesterday
    Auto-check: notes
  • Open Brain Local HTTP

    NateBJones-Projects/OB1

    Captures and searches personal notes in a self-hosted Open Brain using plain curl calls over HTTP, for setups where MCP is disabled or blocked.

    4.7k GitHub stars~1.2k tokensUpdated yesterday
    Auto-check: notes
  • Panning for Gold

    NateBJones-Projects/OB1

    Processes voice transcripts and brain dumps in three phases, extracting every idea thread, evaluating the strongest and saving a permanent inventory and synthesis.

    4.7k GitHub stars~5k tokensUpdated yesterday
    Auto-check passed
  • Agentic Harness Design and Review

    NateBJones-Projects/OB1

    Designs, evaluates and improves the harness around an AI agent: tool permissions, approval gates, state, memory, evals and observability, with phased plans.

    4.7k GitHub stars~1.8k tokensUpdated yesterday
    Auto-check passed

Questions about Heavy File Ingestion Claude Code

What does Heavy File Ingestion Claude Code do?

Use in Claude Code when a user asks to read, analyze, summarize, or extract from a heavyweight file such as PDF, DOCX, PPTX, XLSX, CSV, or TSV. Heavy File Ingestion Claude Code is an agent skill from NateBJones-Projects/OB1. Use in Claude Code when a user asks to read, analyze, summarize, or extract from a heavyweight file such as PDF, DOCX, PPTX, XLSX, CSV, or TSV.

When should I use Heavy File Ingestion Claude Code?

Heavy File Ingestion Claude Code fits situations like: tasks that involve CSV and tabular files; tasks that involve Excel spreadsheets; tasks that involve Word documents.

How do I install Heavy File Ingestion Claude Code in Claude Code?

Run `npx skills add NateBJones-Projects/OB1 --skill heavy-file-ingestion-claude-code -a claude-code`. Or copy the skill folder (skills/heavy-file-ingestion/variants/claude-code in NateBJones-Projects/OB1) into .claude/skills/heavy-file-ingestion-claude-code in your project. Claude Code loads it when a task matches its description.

How do I install Heavy File Ingestion Claude Code in Codex?

Run `npx skills add NateBJones-Projects/OB1 --skill heavy-file-ingestion-claude-code -a codex`. Or copy the skill folder (skills/heavy-file-ingestion/variants/claude-code in NateBJones-Projects/OB1) into .agents/skills/heavy-file-ingestion-claude-code in your project. Codex loads it when a task matches its description.

Can I use Heavy File Ingestion Claude Code in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NateBJones-Projects/OB1 --skill heavy-file-ingestion-claude-code -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/heavy-file-ingestion-claude-code, .gemini/skills/heavy-file-ingestion-claude-code, .github/skills/heavy-file-ingestion-claude-code and .opencode/skills/heavy-file-ingestion-claude-code in your project.

What does Heavy File Ingestion Claude Code need to run?

Going by SKILL.md and its folder, Heavy File Ingestion Claude Code needs the command-line tools its instructions call (python and uv). Our summary lists: Python 3.

Does Heavy File Ingestion Claude Code access the network?

SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Heavy File Ingestion Claude Code safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Heavy File Ingestion Claude Code use?

Heavy File Ingestion Claude Code has a licence file (the repository's licence) that doesn't match a standard licence. Read it on GitHub before reusing the skill.

How many tokens does Heavy File Ingestion Claude Code use?

About 614 tokens (SKILL.md is roughly 2.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Heavy File Ingestion Claude Code?

Skills that share tags, products or a category with Heavy File Ingestion Claude Code: Light File Reading (Light0305/Light-skills, 640 stars), Document Converter (wentorai/Research-Claw, 858 stars), Doc Summarizer (BlackBeltTechnology/pi-agent-dashboard, 315 stars) and To Markdown (Mathews-Tom/armory, 329 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Heavy File Ingestion Claude Code?

NateBJones-Projects (a GitHub organization) maintains it in NateBJones-Projects/OB1, which has 4,710 GitHub stars. The repository holds 24 skills in this directory. The repository was last updated on October 9, 2026.

Source: NateBJones-Projects/OB1 on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.