Discover ML
rand/cc-polymath
Automatically discover machine learning and AI skills when working with machine learning, PyTorch, training, inference, RAG, embeddings, fine-tuning, LLM, DSPy, HuggingFace, or diffusion models.
Methodology for systematically designing document chunking strategies for RAG pipelines.
$ npx skills add revfactory/harness-100 --skill chunking-strategy-guide -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install revfactory/harness-100 chunking-strategy-guide --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/revfactory/harness-100.git skills-src && mkdir -p .claude/skills && cp -r skills-src/en/41-llm-app-builder/.claude/skills/chunking-strategy-guide .claude/skills/chunking-strategy-guide && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "chunking-strategy-guide" agent skill from https://github.com/revfactory/harness-100/tree/main/en/41-llm-app-builder/.claude/skills/chunking-strategy-guide into .claude/skills/chunking-strategy-guide/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "chunking-strategy-guide", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/revfactory/harness-100/tree/main/en/41-llm-app-builder/.claude/skills/chunking-strategy-guideType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add revfactory/harness-100 --skill chunking-strategy-guide -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install revfactory/harness-100 chunking-strategy-guide --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/revfactory/harness-100.git skills-src && mkdir -p .agents/skills && cp -r skills-src/en/41-llm-app-builder/.claude/skills/chunking-strategy-guide .agents/skills/chunking-strategy-guide && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "chunking-strategy-guide" agent skill from https://github.com/revfactory/harness-100/tree/main/en/41-llm-app-builder/.claude/skills/chunking-strategy-guide into .agents/skills/chunking-strategy-guide/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "chunking-strategy-guide", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add revfactory/harness-100 --skill chunking-strategy-guide -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install revfactory/harness-100 chunking-strategy-guide --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/revfactory/harness-100.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/en/41-llm-app-builder/.claude/skills/chunking-strategy-guide .cursor/skills/chunking-strategy-guide && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "chunking-strategy-guide" agent skill from https://github.com/revfactory/harness-100/tree/main/en/41-llm-app-builder/.claude/skills/chunking-strategy-guide into .cursor/skills/chunking-strategy-guide/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "chunking-strategy-guide", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/revfactory/harness-100.git --path en/41-llm-app-builder/.claude/skills/chunking-strategy-guide--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add revfactory/harness-100 --skill chunking-strategy-guide -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install revfactory/harness-100 chunking-strategy-guide --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/revfactory/harness-100.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/en/41-llm-app-builder/.claude/skills/chunking-strategy-guide .gemini/skills/chunking-strategy-guide && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "chunking-strategy-guide" agent skill from https://github.com/revfactory/harness-100/tree/main/en/41-llm-app-builder/.claude/skills/chunking-strategy-guide into .gemini/skills/chunking-strategy-guide/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "chunking-strategy-guide", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install revfactory/harness-100 chunking-strategy-guideInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add revfactory/harness-100 --skill chunking-strategy-guide -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/revfactory/harness-100.git skills-src && mkdir -p .github/skills && cp -r skills-src/en/41-llm-app-builder/.claude/skills/chunking-strategy-guide .github/skills/chunking-strategy-guide && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "chunking-strategy-guide" agent skill from https://github.com/revfactory/harness-100/tree/main/en/41-llm-app-builder/.claude/skills/chunking-strategy-guide into .github/skills/chunking-strategy-guide/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "chunking-strategy-guide", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add revfactory/harness-100 --skill chunking-strategy-guide -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install revfactory/harness-100 chunking-strategy-guide --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/revfactory/harness-100.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/en/41-llm-app-builder/.claude/skills/chunking-strategy-guide .opencode/skills/chunking-strategy-guide && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "chunking-strategy-guide" agent skill from https://github.com/revfactory/harness-100/tree/main/en/41-llm-app-builder/.claude/skills/chunking-strategy-guide into .opencode/skills/chunking-strategy-guide/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "chunking-strategy-guide", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
chunking-strategy-guideMethodology for systematically designing document chunking strategies for RAG pipelines.
Chunking Strategy Guide is an agent skill from revfactory/harness-100. Methodology for systematically designing document chunking strategies for RAG pipelines. Use this skill for 'chunking strategy', 'document splitting', 'RAG chunking', 'embedding optimization', 'semantic chunking', 'text splitting', and other RAG data preprocessing tasks. Note: vector DB infrastructure construction and embedding model training are outside the scope of this skill.
Its SKILL.md is about 1.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in AI & LLM Engineering, covering Embeddings, Retrieval-augmented generation and Machine learning. The licence is Apache-2.0.
Read from SKILL.md and the folder at commit 8e8d35c. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are python).
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Chunking Strategy Guide loads about 1.3k tokens when it runs. Until then it costs about 101 tokens; SKILL.md has 216 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from revfactory/harness-100 at commit 8e8d35c, republished under its Apache-2.0 licence (© revfactory). 216 words, ~1,268 tokens.
.claude/skills/chunking-strategy-guide/SKILL.md (or your agent's skills folder).A skill that enhances the data preprocessing capabilities of the rag-architect.
| Strategy | Principle | Advantages | Disadvantages | Best For |
|---|---|---|---|---|
| Fixed-size | Cut at N-token intervals | Simple implementation | Semantic breaks | Logs, code |
| Sentence-based | Split by sentence | Preserves meaning | Uneven sizes | News, blogs |
| Paragraph-based | Split at blank lines | Maintains logical units | Paragraph size variance | Documents, reports |
| Semantic | Based on embedding similarity | Highest quality | Slow, costly | Complex documents |
| Recursive | Hierarchical separators | Balanced | Complex configuration | General-purpose |
| Markdown | Based on headings | Preserves structure | MD-only | Technical docs |
| Document Type | Chunk Size | Overlap | Rationale |
|--------------|-----------|---------|-----------|
| FAQ | 100-200 tokens | 0 | Q&A pairs are short |
| Technical docs | 300-500 tokens | 50 | Code + explanation units |
| Legal documents | 500-800 tokens | 100 | Article units |
| Academic papers | 400-600 tokens | 80 | Paragraph units |
| Chat logs | 200-300 tokens | 30 | Conversation turn units |
| Fiction/essays | 300-500 tokens | 50 | Scene/paragraph units |optimal_overlap = chunk_size * 0.1 to 0.2
Rules:
- Independent documents (FAQ): Overlap 0
- Sequential documents (manuals): 10-15%
- Dense documents (legal): 15-20%
- Maximum overlap: Never exceed 25% of chunk_sizedef semantic_chunking(text, model, threshold=0.5):
"""
1. Split into sentences
2. Calculate cosine similarity of adjacent sentence embeddings
3. Split at points where similarity falls below threshold
4. Apply min/max chunk size constraints
"""
sentences = split_sentences(text)
embeddings = model.encode(sentences)
breakpoints = []
for i in range(len(embeddings) - 1):
sim = cosine_similarity(embeddings[i], embeddings[i+1])
if sim < threshold:
breakpoints.append(i + 1)
chunks = split_at(sentences, breakpoints)
return enforce_size_limits(chunks, min=100, max=800)PDF > Text extraction (pdfplumber/pymupdf)
> Remove headers/footers
> Remove page numbers
> Tables > Markdown conversion
> Images > alt text / OCR
> Metadata extraction (title, author, date)
> ChunkingHTML > Body extraction (trafilatura/readability)
> Remove navigation/sidebar/ads
> Markdown conversion
> Preserve link text ([text](URL))
> Preserve tables
> Metadata extraction (title, description)
> ChunkingCode > AST parsing
> Split by function/class
> Docstring + signature + body
> Add file path metadata
> Link related test code
> Chunking (function-level)chunk_with_metadata = {
"text": "Chunk text...",
"metadata": {
"source": "document.pdf",
"page": 5,
"section": "3.2 Architecture",
"heading_hierarchy": ["3. Design", "3.2 Architecture"],
"chunk_index": 12,
"total_chunks": 45,
"created_at": "2025-01-15",
"document_type": "technical_spec",
"language": "en"
}
}| Metric | Formula | Threshold |
|---|---|---|
| Information completeness | Original key info / total key info | >= 95% |
| Semantic break rate | Mid-sentence cuts / total chunks | <= 5% |
| Size uniformity | 1 - (std / mean) | >= 0.7 |
| Retrieval precision | Relevant chunks / returned chunks (top-5) | >= 60% |
| Retrieval recall | Returned relevant / total relevant (top-10) | >= 80% |
| Model | Dimensions | Multilingual | Cost | Use Case |
|---|---|---|---|---|
| text-embedding-3-small | 1536 | Good | $0.02/1M | General-purpose, cost-efficient |
| text-embedding-3-large | 3072 | Good | $0.13/1M | High quality |
| multilingual-e5-large | 1024 | Excellent | Free (local) | Multilingual specialization |
| bge-m3 | 1024 | Excellent | Free (local) | Multilingual, long context |
| voyage-multilingual-2 | 1024 | Excellent | $0.12/1M | Best multilingual |
© revfactory, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in en/41-llm-app-builder/.claude/skills/chunking-strategy-guide of revfactory/harness-100.
Open the folder on GitHubat commit 8e8d35c
Chunking Strategy Guide next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Chunking Strategy Guide this skillrevfactory/harness-100 | 1.3k | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | |
| Discover MLrand/cc-polymath | 181 | 1 repos | ~574 | Automated safety check: Pass | MIT | |
| Chroma Vector DatabaseOrchestra-Research/AI-Research-SKILLs | 13k | 8 repos | ~2.3k | Automated safety check: Pass | MIT | |
| Ms Agent Framework RAGshuyu-labs/WebCode | 278 | — | ~1.1k | Automated safety check: Pass | Custom licence | |
| Pgvector Semantic Searchtimescale/pg-aiguide | 1.9k | 1 repos | ~3.8k | Automated safety check: Pass | Apache-2.0 | |
| Evaluate RAGai-evals-course/evals-skills | 1.5k | — | ~1.9k | Automated safety check: Pass | Apache-2.0 |
rand/cc-polymath
Automatically discover machine learning and AI skills when working with machine learning, PyTorch, training, inference, RAG, embeddings, fine-tuning, LLM, DSPy, HuggingFace, or diffusion models.
Orchestra-Research/AI-Research-SKILLs
Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.
shuyu-labs/WebCode
Comprehensive guide for building Agentic RAG systems using Microsoft Agent Framework in C.
timescale/pg-aiguide
A skill your agent uses for setting up vector similarity search with pgvector for AI/ML embeddings, RAG applications, or semantic search.
ai-evals-course/evals-skills
Guides evaluation of a RAG system by diagnosing failures in traces, building a retrieval test set and scoring retrieval and generation separately.
Jeffallan/claude-skills
Designs retrieval-augmented generation systems: document chunking, embeddings, vector store setup, hybrid search, reranking and retrieval evaluation, with checks at each step.
revfactory/harness-100
A skill for analyzing website anti-bot defense mechanisms and developing legitimate evasion strategies.
revfactory/harness-100
Reference for designing how an API reports failures: structured error codes, response shapes, client-friendly messages, an error catalog and retry or fallback advice.
revfactory/harness-100
Walks a backend-dev agent through OWASP API Top 10 checks, authentication and authorization patterns, and defense code during API design.
revfactory/harness-100
Methodology for systematically designing and generating CLI tool argument parser structures.
revfactory/harness-100
Audience segmentation skill used by the analyst and curator agents.
revfactory/harness-100
Audio storytelling skill used by the podcast scriptwriter and show note editor.
Categories
Methodology for systematically designing document chunking strategies for RAG pipelines. Chunking Strategy Guide is an agent skill from revfactory/harness-100. Methodology for systematically designing document chunking strategies for RAG pipelines.
Chunking Strategy Guide fits situations like: chunking strategy; document splitting; embedding optimization; semantic chunking.
Run `npx skills add revfactory/harness-100 --skill chunking-strategy-guide -a claude-code`. Or copy the skill folder (en/41-llm-app-builder/.claude/skills/chunking-strategy-guide in revfactory/harness-100) into .claude/skills/chunking-strategy-guide in your project. Claude Code loads it when a task matches its description.
Run `npx skills add revfactory/harness-100 --skill chunking-strategy-guide -a codex`. Or copy the skill folder (en/41-llm-app-builder/.claude/skills/chunking-strategy-guide in revfactory/harness-100) into .agents/skills/chunking-strategy-guide in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add revfactory/harness-100 --skill chunking-strategy-guide -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/chunking-strategy-guide, .gemini/skills/chunking-strategy-guide, .github/skills/chunking-strategy-guide and .opencode/skills/chunking-strategy-guide in your project.
SKILL.md names no scripts, command-line tools or credentials: Chunking Strategy Guide is instructions for the agent only. Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Chunking Strategy Guide is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.3k tokens (SKILL.md is roughly 5.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Chunking Strategy Guide: Discover ML (rand/cc-polymath, 181 stars), Chroma Vector Database (Orchestra-Research/AI-Research-SKILLs, 13k stars), Ms Agent Framework RAG (shuyu-labs/WebCode, 278 stars) and Pgvector Semantic Search (timescale/pg-aiguide, 1.9k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
revfactory (a GitHub user) maintains it in revfactory/harness-100, which has 1,293 GitHub stars. The repository holds 96 skills in this directory. The repository was last updated on March 22, 2026.
Source: revfactory/harness-100 on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.