A skill your agent uses when splitting extracted text into chunks for LLM context windows or RAG ingestion.

Apache-2.0Auto-check passedAgent Workflows

Install Chunking

skills CLI
$ npx skills add hashgraph-online/awesome-codex-plugins --skill chunking -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install hashgraph-online/awesome-codex-plugins chunking --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/hashgraph-online/awesome-codex-plugins.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/kreuzberg-dev/plugins/plugins/kreuzberg/skills/chunking .claude/skills/chunking && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
chunking
GitHub stars
1.2k
Token cost
~1.4k tokens
SKILL.md length
463 words
Files
1
Skills in repo
686
Repo updated
First seen
Licence
Apache-2.0

At a glance

A skill your agent uses when splitting extracted text into chunks for LLM context windows or RAG ingestion.

  • Splitting extracted text into chunks for LLM context windows
  • SKILL.md covers Inline during extraction, Standalone chunk command, Chunker types and Token-based sizing, plus 4 more sections
  • Calls jq
  • Tasks that involve Retrieval-augmented generation

What it does

Chunking is an agent skill from hashgraph-online/awesome-codex-plugins. Use when splitting extracted text into chunks for LLM context windows or RAG ingestion. Covers chunk size, overlap, markdown/yaml/semantic chunkers, tokenizer-based sizing, and the standalone chunk command.

Its SKILL.md is about 1.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Agent Workflows, covering Retrieval-augmented generation and Context engineering. The repository describes itself as: A curated list of awesome OpenAI Codex / ChatGPT plugins, skills, and resources. The 1 Codex Marketplace. See live plugins at: https://hol.org/plugins/best-codex-plugins. The licence is Apache-2.0.

When your agent uses it

  • Splitting extracted text into chunks for LLM context windows
  • Tasks that involve Retrieval-augmented generation
  • Tasks that involve Context engineering

Example prompts

  • “/chunking”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 78497e5. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • jq

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Chunking loads about 1.4k tokens when it runs. Until then it costs about 54 tokens; SKILL.md has 463 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~54
When it runs · the whole SKILL.md, loaded when a task matches
~1.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from hashgraph-online/awesome-codex-plugins at commit 78497e5, republished under its Apache-2.0 licence (© hashgraph-online). 463 words, ~1,409 tokens.

Download SKILL.mdSave it as .claude/skills/chunking/SKILL.md (or your agent's skills folder).
name
chunking
description
Use when splitting extracted text into chunks for LLM context windows or RAG ingestion. Covers chunk size, overlap, markdown/yaml/semantic chunkers, tokenizer-based sizing, and the standalone `chunk` command.

Chunking

Use this when feeding documents into an LLM context window or a vector store. Kreuzberg chunks two ways: inline during extraction (chunks land on result.chunks), or standalone via the chunk command for text you already have. Sizing is character-based by default, or token-based when a tokenizer model is supplied.

Inline during extraction

Turn on chunking with --chunk and the chunks appear on the structured result under chunks:

bash
# 1000-char chunks, 200-char overlap (defaults when --chunk is on)
kreuzberg extract report.pdf --chunk --format json | jq '.chunks | length'

# Explicit size + overlap
kreuzberg extract report.pdf --chunk --chunk-size 1500 --chunk-overlap 300 --format json

Overlap must be smaller than chunk size — the CLI rejects --chunk-overlap >= --chunk-size. When you set only --chunk-overlap against an existing config, an overlap that exceeds the size is clamped to chunk_size / 4.

Standalone chunk command

Chunk text you already have, from --text or stdin. Output defaults to JSON:

bash
# From a flag
kreuzberg chunk --text "long document text ..." --chunk-size 800 --chunk-overlap 100

# From stdin (pipe extracted content straight in)
kreuzberg extract notes.md | kreuzberg chunk --chunk-size 500 --format json

JSON output carries chunks (array of strings), chunk_count, the resolved config (max_characters, overlap, chunker_type), and input_size_bytes. Use --format text for a human-readable dump with --- chunk N --- separators.

Note: in the JSON output, chunker_type is rendered capitalized ("Text", "Markdown", "Yaml", "Semantic") because it is emitted via Rust's Debug formatting, whereas the --chunker-type input flag is lowercase (text, markdown, yaml, semantic). Lowercase the value before comparing if you parse it back.

Chunker types

--chunker-type selects the splitting strategy (standalone chunk command):

TypeBehavior
textDefault. Plain character-window splitting with overlap.
markdownMarkdown-aware — splits on structure (headings, blocks) where possible.
yamlYAML-aware splitting for structured config/data documents.
semanticTopic-boundary splitting driven by --topic-threshold (0.0–1.0, default 0.75).
bash
# Markdown-aware chunking keeps headings and blocks intact
kreuzberg chunk --text "$(cat README.md)" --chunker-type markdown

# Semantic chunking — lower threshold = more, smaller topic chunks
kreuzberg chunk --text "$(cat transcript.txt)" --chunker-type semantic --topic-threshold 0.6

Token-based sizing

By default --chunk-size counts characters. To size chunks by tokens for a specific model, pass --chunking-tokenizer with a HuggingFace tokenizer id. On the extract command this implicitly enables chunking. Requires the chunking-tokenizers feature (present in the default CLI build).

bash
# Size chunks by GPT-4o tokens during extraction
kreuzberg extract report.pdf --chunking-tokenizer Xenova/gpt-4o --format json

# Or on the standalone command
kreuzberg chunk --text "$(cat doc.txt)" --chunking-tokenizer Xenova/gpt-4o --chunk-size 512

With a tokenizer set, --chunk-size is interpreted in tokens, not characters.

Show full SKILL.md (182 more words)Show less

Config file alternative

Field names in config files are snake_case under [chunking]:

toml
[chunking]
max_characters = 1000
overlap = 200
chunker_type = "markdown"
bash
kreuzberg extract report.pdf --config kreuzberg.toml --format json

CLI flags map to config fields as --chunk-size → max_characters and --chunk-overlap → overlap. In config files use the snake_case names.

Programmatic access

From Python, enable chunking on the config and read result.chunks:

python
from kreuzberg import extract_file_sync, ExtractionConfig, ChunkingConfig

config = ExtractionConfig(
    chunking=ChunkingConfig(max_chars=1000, max_overlap=200),
)
result = extract_file_sync("report.pdf", config=config)
for chunk in result.chunks:
    print(len(chunk))

Python ChunkingConfig uses max_chars / max_overlap. Rust uses max_characters / overlap. See references/python-api.md and references/rust-api.md in the sibling kreuzberg skill.

Picking parameters

  • RAG / vector store — 500–1000 chars (or 256–512 tokens) with 10–20% overlap. Use markdown chunking for docs to keep sections whole.
  • LLM summarization — larger chunks (1500–4000 chars) with small overlap; size by tokens to stay under the model window.
  • Topic segmentation — semantic chunker; tune --topic-threshold down for finer splits, up for coarser ones.

Common pitfalls

  • Overlap ≥ size — rejected on extract; clamped to size / 4 when only overlap is changed against an existing config.
  • Tokenizer without the feature — --chunking-tokenizer errors if the CLI was built without chunking-tokenizers. The default build includes it.
  • Empty input — the standalone chunk command bails on empty text; provide --text or pipe non-empty stdin.

See references/configuration.md for the full [chunking] schema and references/cli-reference.md for every chunk flag.

© hashgraph-online, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in plugins/kreuzberg-dev/plugins/plugins/kreuzberg/skills/chunking of hashgraph-online/awesome-codex-plugins.

Open the folder on GitHubat commit 78497e5

Compare with similar skills

Chunking next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Chunking compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Chunking this skillhashgraph-online/awesome-codex-plugins1.2k—~1.4kAutomated safety check: PassApache-2.0
Context Engineborghei/Claude-Skills881—~2.2kAutomated safety check: PassMIT
Context Retrievalseb1n/awesome-ai-agent-skills206—~2.1kAutomated safety check: PassMIT
Memori Long-Term MemoryMemoriLabs/Memori17k—~2kAutomated safety check: PassApache-2.0
Project Developmentguanyang/open-agent-hub9752 repos~4.7kAutomated safety check: PassMIT
Context DoctorjzOcb/context-doctor119—~642Automated safety check: PassMIT

Similar skills

  • Context Engine

    borghei/Claude-Skills

    Context management engine for AI coding agents. An agent skill from borghei/Claude-Skills.

    881 GitHub stars~2.2k tokensUpdated yesterday
    Agent WorkflowsAuto-check passed
  • Context Retrieval

    seb1n/awesome-ai-agent-skills

    Retrieve relevant information from a knowledge base using semantic, keyword, or hybrid search to ground a query.

    206 GitHub stars~2.1k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed
  • Memori Long-Term Memory

    MemoriLabs/Memori

    Adds structured long-term memory to OpenClaw agents, built automatically from sessions, with tools the agent calls to recall facts, summaries and decisions.

    17k GitHub stars~2k tokensUpdated 5 days ago
    Agent WorkflowsAuto-check passed
  • Project Development

    guanyang/open-agent-hub

    This skill should be used for project-level decisions about LLM-powered systems: whether an LLM is the right primitive for the task at hand, the shape of a multi-stage batch or agent pipeline, token…

    975 GitHub starsUsed in 2 repos~4.7k tokens
    Agent WorkflowsAuto-check passed
  • Context Doctor

    jzOcb/context-doctor

    Visualize and diagnose OpenClaw context window usage. An agent skill from jzOcb/context-doctor.

    119 GitHub stars~642 tokensUpdated 6 mo ago
    Agent WorkflowsAuto-check passed
  • Clawmem

    yoloshii/ClawMem

    ClawMem operational reference for agents at query time — the 3-rule escalation gate, MCP tool routing, the 4 query-optimization levers, pipeline behavior (query vs intentsearch), composite scoring…

    210 GitHub stars~7.5k tokensUpdated 3 days ago
    Agent WorkflowsAuto-check passed

More from hashgraph-online/awesome-codex-plugins

All 686 skills in this repo
  • Anime Reaction Gif

    hashgraph-online/awesome-codex-plugins

    Create original anime-style reaction stickers as looping GIFs and MP4 previews, using generated character pose sheets and timed key poses.

    1.2k GitHub stars~922 tokensUpdated today
    Auto-check passed
  • Calibredb

    hashgraph-online/awesome-codex-plugins

    Manage and query Calibre libraries with the calibredb CLI (local paths or Calibre Content server URLs).

    1.2k GitHub stars~1k tokensUpdated today
    Auto-check passed
  • Rust API Test Harness

    hashgraph-online/awesome-codex-plugins

    A skill your agent uses when adding, changing, testing, or debugging Rust HTTP APIs and services, especially when Codex needs black-box integration tests, random-port app startup, real database test…

    1.2k GitHub stars~1.7k tokensUpdated today
    Auto-check passed
  • Art

    hashgraph-online/awesome-codex-plugins

    Make a studio's game look like something at build time — a cover from a real frame of the game (free), painted covers, backdrops, textures and character plates from image models through the…

    1.2k GitHub stars~2.4k tokensUpdated today
    Auto-check passed
  • Game Balance Economy

    hashgraph-online/awesome-codex-plugins

    Balance game difficulty, resources, rewards, probability, progression, economies, and dominant strategies.

    1.2k GitHub stars~618 tokensUpdated today
    Auto-check passed
  • Manuscript Engagement Analytics

    hashgraph-online/awesome-codex-plugins

    Analyze nonfiction manuscripts for reader engagement signals, including heading-level word counts, slow starts, long slogs, weak takeaway titles, value pacing, beta-reader comment dropoff, and…

    1.2k GitHub stars~875 tokensUpdated today
    Auto-check passed

Questions about Chunking

What does Chunking do?

A skill your agent uses when splitting extracted text into chunks for LLM context windows or RAG ingestion. Chunking is an agent skill from hashgraph-online/awesome-codex-plugins. Use when splitting extracted text into chunks for LLM context windows or RAG ingestion.

When should I use Chunking?

Chunking fits situations like: splitting extracted text into chunks for LLM context windows; tasks that involve Retrieval-augmented generation; tasks that involve Context engineering.

How do I install Chunking in Claude Code?

Run `npx skills add hashgraph-online/awesome-codex-plugins --skill chunking -a claude-code`. Or copy the skill folder (plugins/kreuzberg-dev/plugins/plugins/kreuzberg/skills/chunking in hashgraph-online/awesome-codex-plugins) into .claude/skills/chunking in your project. Claude Code loads it when a task matches its description.

How do I install Chunking in Codex?

Run `npx skills add hashgraph-online/awesome-codex-plugins --skill chunking -a codex`. Or copy the skill folder (plugins/kreuzberg-dev/plugins/plugins/kreuzberg/skills/chunking in hashgraph-online/awesome-codex-plugins) into .agents/skills/chunking in your project. Codex loads it when a task matches its description.

Can I use Chunking in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add hashgraph-online/awesome-codex-plugins --skill chunking -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/chunking, .gemini/skills/chunking, .github/skills/chunking and .opencode/skills/chunking in your project.

What does Chunking need to run?

Going by SKILL.md and its folder, Chunking needs the command-line tools its instructions call (jq). Our summary lists: Python 3.

Does Chunking access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Chunking safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Chunking use?

Chunking is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Chunking use?

About 1.4k tokens (SKILL.md is roughly 5.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Chunking?

Skills that share tags, products or a category with Chunking: Context Engine (borghei/Claude-Skills, 881 stars), Context Retrieval (seb1n/awesome-ai-agent-skills, 206 stars), Memori Long-Term Memory (MemoriLabs/Memori, 17k stars) and Project Development (guanyang/open-agent-hub, 975 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Chunking?

hashgraph-online (a GitHub organization) maintains it in hashgraph-online/awesome-codex-plugins, which has 1,242 GitHub stars. The repository holds 686 skills in this directory. The repository was last updated on October 8, 2026.

Source: hashgraph-online/awesome-codex-plugins on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.