A skill your agent uses when extracting from many files at once with shared config, bounded parallelism, per-file overrides, and error recovery.

Apache-2.0Auto-check passed

Install Batch Extraction

skills CLI
$ npx skills add hashgraph-online/awesome-codex-plugins --skill batch-extraction -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install hashgraph-online/awesome-codex-plugins batch-extraction --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/hashgraph-online/awesome-codex-plugins.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/kreuzberg-dev/plugins/plugins/kreuzberg/skills/batch-extraction .claude/skills/batch-extraction && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
batch-extraction
GitHub stars
1.2k
Token cost
~1.3k tokens
SKILL.md length
441 words
Files
1
Skills in repo
736
Repo updated
First seen
Licence
Apache-2.0

At a glance

A skill your agent uses when extracting from many files at once with shared config, bounded parallelism, per-file overrides, and error recovery.

  • Extracting from many files at once with shared config
  • SKILL.md covers Basic usage, Parallelism, Per-file config overrides and Output layout, plus 5 more sections
  • Calls jq
  • Bounded parallelism

What it does

Batch Extraction is an agent skill from hashgraph-online/awesome-codex-plugins. Use when extracting from many files at once with shared config, bounded parallelism, per-file overrides, and error recovery. Covers the batch command, --file-configs, --max-concurrent, and output layout.

Its SKILL.md is about 1.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

The repository describes itself as: A curated list of awesome OpenAI Codex / ChatGPT plugins, skills, and resources. The 1 Codex Marketplace. See live plugins at: https://hol.org/plugins/best-codex-plugins. The licence is Apache-2.0.

When your agent uses it

  • Extracting from many files at once with shared config
  • Bounded parallelism
  • Per-file overrides

Example prompts

  • “/batch-extraction”

Requirements

  • Python 3
  • Node.js

What it can do on your machine

Read from SKILL.md and the folder at commit 16b4156. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • jq

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Batch Extraction loads about 1.3k tokens when it runs. Until then it costs about 57 tokens; SKILL.md has 441 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~57
When it runs · the whole SKILL.md, loaded when a task matches
~1.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from hashgraph-online/awesome-codex-plugins at commit 16b4156, republished under its Apache-2.0 licence (© hashgraph-online). 441 words, ~1,250 tokens.

Download SKILL.mdSave it as .claude/skills/batch-extraction/SKILL.md (or your agent's skills folder).
name
batch-extraction
description
Use when extracting from many files at once with shared config, bounded parallelism, per-file overrides, and error recovery. Covers the `batch` command, `--file-configs`, `--max-concurrent`, and output layout.

Batch extraction

Use this when processing a directory or glob of documents in one pass. kreuzberg batch shares one extraction config across every file, runs extractions concurrently, and returns one structured array — failures on individual files do not abort the run.

Basic usage

bash
# Glob expands to many paths; results come back as a JSON array (default)
kreuzberg batch *.pdf

# Mixed formats, markdown content for LLM ingestion
kreuzberg batch docs/*.docx --content-format markdown

# Recurse with the shell, then extract
kreuzberg batch $(find ./corpus -name '*.pdf')

batch defaults to --format json (vs --format text for single extract). Each array entry is a full extraction result, so downstream code can index by position into the input path list.

bash
kreuzberg batch reports/*.pdf \
  | jq '.[] | {chars: (.content | length), mime: .mime_type}'

Parallelism

--max-concurrent caps how many files extract at once (default: CPU count). Lower it on memory-constrained hosts or when OCR/ML models are active, since each in-flight extraction holds its own buffers:

bash
# Cap at 4 concurrent extractions
kreuzberg batch scans/*.pdf --ocr true --max-concurrent 4

--max-threads additionally caps total internal threads (Rayon, ONNX intra-op, the batch semaphore) for tightly constrained environments:

bash
kreuzberg batch *.pdf --max-concurrent 2 --max-threads 4

Per-file config overrides

A single shared config does not always fit. --file-configs points at a JSON file mapping each path to its own override object, merged on top of the shared config for that file only:

json
{
  "scan.pdf": { "force_ocr": true },
  "report.pdf": { "output_format": "markdown" },
  "data.xlsx": { "output_format": "json" }
}
bash
kreuzberg batch scan.pdf report.pdf data.xlsx --file-configs overrides.json

Keys are file paths (matching the paths passed on the command line); values are per-file extraction config objects in snake_case, the same shape as a config file.

Output layout

For text/toon output with image extraction, --output-dir controls where referenced image files (e.g. image_0.png) are written; the directory must already exist. JSON output embeds image bytes inline and ignores --output-dir.

bash
mkdir -p out/images
kreuzberg batch slides/*.pptx --extract-images true --output-dir out/images --format text

Error recovery

Batch extraction is fault-tolerant per file: one unreadable or corrupt document does not stop the rest. Inspect results for partial content and surfaced errors rather than relying on the process exit code alone. Pair with --max-concurrent to avoid exhausting memory when a few large files sit in a big batch.

Show full SKILL.md (173 more words)Show less

Shared config

Every extract flag also applies to batch (OCR, chunking, layout, content format, etc.) and is shared across all files unless a --file-configs entry overrides it:

bash
kreuzberg batch invoices/*.pdf \
  --layout --layout-table-model slanet_wireless \
  --content-format markdown --max-concurrent 8

A config file works too and auto-discovers from the cwd upward:

toml
output_format = "markdown"

[ocr]
backend = "tesseract"
language = "eng"
bash
kreuzberg batch corpus/*.pdf --config kreuzberg.toml

Programmatic access

From Python, use the batch helpers (async and sync):

python
from kreuzberg import batch_extract_files, batch_extract_files_sync, ExtractionConfig

config = ExtractionConfig(output_format="markdown")

# Async
results = await batch_extract_files(["a.pdf", "b.docx", "c.xlsx"], config=config)

# Sync
results = batch_extract_files_sync(["a.pdf", "b.docx"], config=config)

for result in results:
    print(len(result.content))

Node.js mirrors this with batchExtractFiles; Rust uses batch_extract_file (requires the tokio-runtime feature). See references/python-api.md, references/nodejs-api.md, and references/rust-api.md in the sibling kreuzberg skill.

MCP

When the kreuzberg MCP server is registered, prefer the batch_extract_files tool over shelling out — it takes the file list and a config object and returns structured results directly.

Common pitfalls

  • Default format differs — batch defaults to --format json, extract to --format text. Set --format explicitly if a script depends on one shape.
  • --output-dir must exist — the CLI does not create it.
  • Memory blowups — large batches with OCR/layout active need a lower --max-concurrent; the default is CPU count.
  • --file-configs path keys — must match the paths as passed on the command line, not absolute-resolved variants.

See references/cli-reference.md for the full batch flag set.

© hashgraph-online, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in plugins/kreuzberg-dev/plugins/plugins/kreuzberg/skills/batch-extraction of hashgraph-online/awesome-codex-plugins.

Open the folder on GitHubat commit 16b4156

Compare with similar skills

Batch Extraction next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Batch Extraction compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Batch Extraction this skillhashgraph-online/awesome-codex-plugins1.2k—~1.3kAutomated safety check: PassApache-2.0
Batchasgeirtj/system_prompts_leaks69k—~1.3kAutomated safety check: PassCC0-1.0
Batchcodewhale-hq/Codewhale41k—~157Automated safety check: PassMIT
Extractalirezarezvani/claude-skills28k—~1.4kAutomated safety check: PassMIT
Batch API PlannerQwenLM/qwen-code28k—~2.2kAutomated safety check: PassApache-2.0
Brand Extractnexu-io/open-design100k—~3.1kAutomated safety check: PassApache-2.0

Similar skills

  • Batch

    asgeirtj/system_prompts_leaks

    Research and plan a large-scale change, then execute it in parallel across 5–30 isolated worktree agents that each open a PR.

    69k GitHub stars~1.3k tokensUpdated yesterday
    DevelopmentAuto-check passed
  • Batch

    codewhale-hq/Codewhale

    Break a large, parallelizable goal into bounded work units, coordinate existing agent/worktree machinery, integrate, and verify.

    41k GitHub stars~157 tokensUpdated today
    DevelopmentAuto-check passed
  • Extract

    alirezarezvani/claude-skills

    Turn a proven pattern or debugging solution into a standalone reusable skill with SKILL.md, reference docs, and examples.

    28k GitHub stars~1.4k tokensUpdated 1 mo ago
    DevelopmentAuto-check passed
  • Batch API Planner

    QwenLM/qwen-code

    Prepares many-file, single-turn transforms such as translating or rewriting as a plan, then submits it to the asynchronous, half-price DashScope Batch API through the qwen batch CLI.

    28k GitHub stars~2.2k tokensUpdated today
    Productivity & AutomationAuto-check passed
  • Brand Extract

    nexu-io/open-design

    Extract a complete Brand Kit from a live website by driving the in-app browser.

    100k GitHub stars~3.1k tokensUpdated today
    Productivity & AutomationAuto-check passed
  • Runs one operation across many files with parallel worker agents: finds files by glob pattern, splits them into chunks, launches workers and summarizes the results.

    28k GitHub stars~2.3k tokensUpdated today
    DevelopmentAuto-check passed

More from hashgraph-online/awesome-codex-plugins

All 736 skills in this repo
  • Anime Reaction Gif

    hashgraph-online/awesome-codex-plugins

    Create original anime-style reaction stickers as looping GIFs and MP4 previews, using generated character pose sheets and timed key poses.

    1.2k GitHub stars~922 tokensUpdated yesterday
    Auto-check passed
  • Calibredb

    hashgraph-online/awesome-codex-plugins

    Manage and query Calibre libraries with the calibredb CLI (local paths or Calibre Content server URLs).

    1.2k GitHub stars~1k tokensUpdated yesterday
    Auto-check passed
  • Rust API Test Harness

    hashgraph-online/awesome-codex-plugins

    A skill your agent uses when adding, changing, testing, or debugging Rust HTTP APIs and services, especially when Codex needs black-box integration tests, random-port app startup, real database test…

    1.2k GitHub stars~1.7k tokensUpdated yesterday
    Auto-check passed
  • Art

    hashgraph-online/awesome-codex-plugins

    Make a studio's game look like something at build time — a cover from a real frame of the game (free), painted covers, backdrops, textures and character plates from image models through the…

    1.2k GitHub stars~2.4k tokensUpdated yesterday
    Auto-check passed
  • Calle

    hashgraph-online/awesome-codex-plugins

    Use CALL-E from Codex through the calle CLI. An agent skill from hashgraph-online/awesome-codex-plugins.

    1.2k GitHub stars~2.9k tokensUpdated yesterday
    Auto-check passed
  • Game Balance Economy

    hashgraph-online/awesome-codex-plugins

    Balance game difficulty, resources, rewards, probability, progression, economies, and dominant strategies.

    1.2k GitHub stars~618 tokensUpdated yesterday
    Auto-check passed

Questions about Batch Extraction

What does Batch Extraction do?

A skill your agent uses when extracting from many files at once with shared config, bounded parallelism, per-file overrides, and error recovery. Batch Extraction is an agent skill from hashgraph-online/awesome-codex-plugins. Use when extracting from many files at once with shared config, bounded parallelism, per-file overrides, and error recovery.

When should I use Batch Extraction?

Batch Extraction fits situations like: extracting from many files at once with shared config; bounded parallelism; per-file overrides.

How do I install Batch Extraction in Claude Code?

Run `npx skills add hashgraph-online/awesome-codex-plugins --skill batch-extraction -a claude-code`. Or copy the skill folder (plugins/kreuzberg-dev/plugins/plugins/kreuzberg/skills/batch-extraction in hashgraph-online/awesome-codex-plugins) into .claude/skills/batch-extraction in your project. Claude Code loads it when a task matches its description.

How do I install Batch Extraction in Codex?

Run `npx skills add hashgraph-online/awesome-codex-plugins --skill batch-extraction -a codex`. Or copy the skill folder (plugins/kreuzberg-dev/plugins/plugins/kreuzberg/skills/batch-extraction in hashgraph-online/awesome-codex-plugins) into .agents/skills/batch-extraction in your project. Codex loads it when a task matches its description.

Can I use Batch Extraction in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add hashgraph-online/awesome-codex-plugins --skill batch-extraction -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/batch-extraction, .gemini/skills/batch-extraction, .github/skills/batch-extraction and .opencode/skills/batch-extraction in your project.

What does Batch Extraction need to run?

Going by SKILL.md and its folder, Batch Extraction needs the command-line tools its instructions call (jq). Our summary lists: Python 3; Node.js.

Does Batch Extraction access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Batch Extraction safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Batch Extraction use?

Batch Extraction is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Batch Extraction use?

About 1.3k tokens (SKILL.md is roughly 5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Batch Extraction?

Skills that share tags, products or a category with Batch Extraction: Batch (asgeirtj/system_prompts_leaks, 69k stars), Batch (codewhale-hq/Codewhale, 41k stars), Extract (alirezarezvani/claude-skills, 28k stars) and Batch API Planner (QwenLM/qwen-code, 28k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Batch Extraction?

hashgraph-online (a GitHub organization) maintains it in hashgraph-online/awesome-codex-plugins, which has 1,232 GitHub stars. The repository holds 736 skills in this directory. The repository was last updated on October 6, 2026.

Source: hashgraph-online/awesome-codex-plugins on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.