Agent skill

Agentsop Multiscale Chunking

by agentsope in agentsope/SkillAlchemy

Designs multiscale chunking for RAG by embedding small units for retrieval precision and returning larger context for synthesis.

MITAuto-check passedAI & LLM Engineering

Install Agentsop Multiscale Chunking

skills CLI
$ npx skills add agentsope/SkillAlchemy --skill agentsop-multiscale-chunking -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install agentsope/SkillAlchemy agentsop-multiscale-chunking --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/agentsope/SkillAlchemy.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/agentsop-multiscale-chunking .claude/skills/agentsop-multiscale-chunking && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
agentsop-multiscale-chunking
GitHub stars
459
Token cost
~4.9k tokens
SKILL.md length
2,085 words
Files
4 (incl. references)
Skills in repo
45
Repo updated
First seen
Licence
MIT

At a glance

Designs multiscale chunking for RAG by embedding small units for retrieval precision and returning larger context for synthesis.

  • Works in 7 steps: 何时激活 (Activation Rules) → 核心心智模型 (Core Mental Model) → SOP 工作流 (Decision Protocol) → …
  • Fixed-size chunks either lose surrounding context
  • SKILL.md covers 1. 何时激活 (Activation Rules), 2. 核心心智模型 (Core Mental Model), 3. SOP 工作流 (Decision Protocol) and 4. 操作模型 (Operation Models), plus 4 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Agentsop Multiscale Chunking is an agent skill from agentsope/SkillAlchemy. Designs multiscale chunking for RAG by embedding small units for retrieval precision and returning larger context for synthesis. Use when fixed-size chunks either lose surrounding context or dilute relevance in long documents, manuals, filings, or codebases. Covers sentence-window and parent-child or auto-merging strategies, base chunk sizing, and evaluation. Do not use when a single chunk scale already meets retrieval and generation needs.

Its SKILL.md is about 4.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including reference files (for example `README.md`, `intermediate/operation_candidates.json` and `references/R1-source-evidence.md`).

It sits in AI & LLM Engineering, covering Retrieval-augmented generation and Embeddings. It works with LlamaIndex. The repository describes itself as: From thought to skill. From signal to structure. The licence is MIT.

When your agent uses it

  • Fixed-size chunks either lose surrounding context
  • Dilute relevance in long documents
  • A single chunk scale already meets retrieval and generation needs

Example prompts

  • “Use the agentsop-multiscale-chunking skill to design multiscale chunking for RAG by embedding small units for retrieval precision and returning…”
  • “/agentsop-multiscale-chunking”

Requirements

  • Python 3

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. 何时激活 (Activation Rules)
  2. 核心心智模型 (Core Mental Model)
  3. SOP 工作流 (Decision Protocol)
  4. 操作模型 (Operation Models)
  5. 困境决策案例 (Dilemma Cases)
  6. 反模式与边界 (Anti-patterns & Boundaries)
  7. 跨框架对照 (Cross-framework Mapping)

What it can do on your machine

Read from SKILL.md and the folder at commit 6ea799f. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Agentsop Multiscale Chunking loads about 4.9k tokens when it runs, and up to ~6.8k if it reads all its reference files. Until then it costs about 118 tokens; SKILL.md has 2,085 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~118
When it runs · the whole SKILL.md, loaded when a task matches
~4.9k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~6.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from agentsope/SkillAlchemy at commit 6ea799f, republished under its MIT licence (© agentsope). 2,085 words, ~4,878 tokens.

Download SKILL.mdSave it as .claude/skills/agentsop-multiscale-chunking/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
agentsop-multiscale-chunking
description
Designs multiscale chunking for RAG by embedding small units for retrieval precision and returning larger context for synthesis. Use when fixed-size chunks either lose surrounding context or dilute relevance in long documents, manuals, filings, or codebases. Covers sentence-window and parent-child or auto-merging strategies, base chunk sizing, and evaluation. Do not use when a single chunk scale already meets retrieval and generation needs.
version
0.1.0

Multi-scale Chunking · C5 Enhancement Overlay

Overlay on top of [[llamaindex]]. The base skill teaches the 5-layer RAG pipeline and lists DecoupleChunkScope as one optimization knob among many. This overlay zooms in on that single knob and turns it into a standalone recipe: how to resolve the chunk paradox when one chunk size is provably not enough. Third-person analytical view for an agent writing / reviewing RAG ingestion code — not an end-user tutorial.


1. 何时激活 (Activation Rules)

Activate this overlay when all three RAG preconditions hold and the chunk paradox has actually surfaced:

  1. The corpus is long documents — prose manuals, financial filings, legal contracts, research papers, codebases — where a single answer-bearing fact sits inside a larger context that the LLM needs to interpret it.
  2. A chunk-size sweep has stalled: small chunks (128–256) win retrieval precision but the LLM answers from fragments; large chunks (1024–2048) give rich context but recall on specific queries drops because the embedding becomes a "topic average". The official failure-mode checklist documents both poles as separate failures — #2 (wrong chunk from too-small) and #6 (context overflow / dilution from too-large) (cited in [[llamaindex]] R3).
  3. Faithfulness or relevancy is plateauing below target and bumping chunk_size only moves the failure from one pole to the other.

Concrete triggers:

  • "Answers are technically retrieved but the model lacks context to explain them."
  • "I keep retuning chunk_size and it never wins on both faithfulness and recall."
  • A reviewer sees SentenceSplitter(chunk_size=4096) shipped as the fix for "incomplete answers" (this is anti-pattern A1 in [[llamaindex]]).

Do not activate when:

  • The corpus is short/static (<100k tokens) — prompt-stuff with caching; multi-scale chunking is over-engineering (§6).
  • The chunk-size sweep did converge on a single winner (e.g. 1024 for prose) — pin it and stop.
  • Retrieval quality is fine and the bottleneck is orchestration or synthesis.

2. 核心心智模型 (Core Mental Model)

Decouple the embed-unit from the return-unit. Embed small for retrieval precision; return large for generation context.

The naive assumption is that the unit you index is the unit you feed the LLM. That single identity is the source of the paradox: it forces one chunk size to serve two opposing jobs.

NAIVE (one unit, two jobs)          MULTI-SCALE (two units, one job each)
─────────────────────────          ─────────────────────────────────────
       [ chunk ]                    embed unit  →  small  (precision job)
      /         \                          │
 embed it     feed it                   match
 (wants       (wants                       │
  small)       large)                 return unit →  large  (context job)
   ↓             ↓                          ▲
  CONFLICT — pick one,                   expand from
  lose the other                        match → parent / window

Three load-bearing sub-principles:

  1. A Node is a graph node, not a chunk. In LlamaIndex a Node carries relationships (PREV/NEXT/PARENT/CHILD). Those links are exactly what let you store a small node for matching and resolve it to a larger node for return ([[llamaindex]] Principle 2). Multi-scale chunking is "build a chunk-graph", not "split into chunks".

  2. Two geometries of expansion. Once embed ≠ return, you must choose how the small match expands into the large return:

    • Horizontal — return N adjacent sentences around the matched sentence (sentence-window). The expansion is positional.
    • Vertical — return the parent chunk when enough sibling children match (auto-merging / parent-child). The expansion is hierarchical.
  3. Match the geometry to the document, not to taste. Flat narrative prose → horizontal. Documents with real structure (headings, sections, tables of contents) → vertical. This is the central dilemma case (§5.1).

The overlay's promise: this strictly dominates a compromise chunk size when the sweep frontier is non-flat — you no longer average two bad sizes.


3. SOP 工作流 (Decision Protocol)

A three-gate protocol. Do not skip Gate 0 — multi-scale chunking is only justified once a single chunk size has been proven insufficient.

Gate 0 — Pick the base chunk size first (and try to stop here)

Run the canonical sweep from [[llamaindex]] OP-02 before reaching for any multi-scale machinery:

python
from llama_index.core.evaluation import (
    FaithfulnessEvaluator, RelevancyEvaluator,
)
# 1. ~20 eval QA pairs via DatasetGenerator.from_documents(docs)
# 2. sweep:
for cs in (128, 256, 512, 1024, 2048):           # overlap = 0.1–0.2 × cs
    idx = build_index(docs, chunk_size=cs, overlap=int(0.15 * cs))
    record(cs, faithfulness=eval_f(idx), relevancy=eval_r(idx), p95=latency(idx))
  • If a single chunk size dominates both faithfulness and relevancy → pin it and stop. LlamaIndex's own Uber 10-K study peaked at 1024 for prose; code lands at 80–160 tokens (§5.2).
  • If the frontier is non-flat (small wins precision, large wins context, no single winner) → proceed to Gate 1. Do not compromise on a middle size.
Gate 1 — Choose the expansion strategy by document structure
Document shapeStrategyGeometry
Flat prose, no clear sectioningSentence-Windowhorizontal
Clear hierarchy (headings, sections, ToC)Auto-Merging (Hierarchical / parent-child)vertical
Bursty multi-chunk relevance ("this whole section matters")Auto-Mergingvertical
Point-fact needing surrounding paragraphSentence-Windowhorizontal
Unknown structure / lowest setup costStart Sentence-Windowhorizontal

Set the embed-unit small (single sentence, or 128–256-token leaf) and the return-unit large (the window, or the parent/root chunk).

Gate 2 — Wire it and measure the lift
  • Sentence-Window: SentenceWindowNodeParser must be paired with MetadataReplacementPostProcessor — otherwise the metadata-stuffed matched sentence (not the window) reaches the LLM, defeating the entire point (§6 A2).
  • Auto-Merging: HierarchicalNodeParser builds leaf+parent nodes; store all leaves in the docstore and AutoMergingRetriever merges children → parent when ≥ threshold siblings match.
  • Re-run the same eval set from Gate 0. Both patterns should beat naive top-k on faithfulness; if neither does, the bottleneck is elsewhere (revert).
  • Pin the chosen parser + retriever config and the embedding-model version into index metadata, exactly as Gate 0's chunk size would have been pinned.

4. 操作模型 (Operation Models)

Each operation: Trigger / Action / Output / Evidence. These refine [[llamaindex]] OP-02 and OP-05 into executable sub-steps.

MSC-01 — ChunkSizeSweep
  • Trigger: New long-doc corpus; chunk size unknown; before any multi-scale work.
  • Action: Generate ~20 QA pairs; sweep chunk_size ∈ {128,256,512,1024,2048}, overlap = 10–20%; build one VectorStoreIndex per config; record faithfulness + relevancy + p95 latency.
  • Output: Either a pinned single winner, OR a documented non-flat frontier that authorizes multi-scale chunking.
  • Evidence: llamaindex.ai/blog/evaluating-the-ideal-chunk-size-for-a-rag-system-using-llamaindex-6207e5d3fec5; [[llamaindex]] OP-02.
MSC-02 — SentenceWindowSetup (horizontal expansion)
  • Trigger: Flat prose; point-facts that need a surrounding paragraph; lowest setup cost wanted.
  • Action: SentenceWindowNodeParser(window_size=3) to embed single sentences with N-neighbor windows in metadata; at query time apply MetadataReplacementPostProcessor(target_metadata_key="window") so the LLM receives the window, not the lone sentence.
  • Output: Precise sentence-level matching + paragraph-level synthesis context.
  • Evidence: developers.llamaindex.ai SentenceWindow / MetadataReplacement docs; [[llamaindex]] R3 Dilemma 5.
MSC-03 — AutoMergingHierarchy (vertical expansion)
  • Trigger: Documents with clear hierarchy; bursty multi-chunk relevance.
  • Action: HierarchicalNodeParser.from_defaults(chunk_sizes=[2048,512,128]) to build a leaf→parent→root tree; index leaf nodes in a VectorStoreIndex, keep all nodes in a docstore; retrieve with AutoMergingRetriever, which returns the parent once a configured fraction of its children are in the hit set.
  • Output: Leaf-level precision that escalates to section-level context when warranted.
  • Evidence: developers.llamaindex.ai/python/framework/integrations/retrievers/auto_merging_retriever/; [[llamaindex]] OP-05.
MSC-04 — WindowVsMergeChoice
  • Trigger: Multi-scale authorized (MSC-01 non-flat) but strategy undecided.
  • Action: Inspect document structure (Gate 1 table). Flat → MSC-02; structured/bursty → MSC-03; unknown → start MSC-02 (cheaper), escalate to MSC-03 if it underperforms on multi-chunk queries.
  • Output: A strategy chosen by document shape, not by elegance.
  • Evidence: [[llamaindex]] R3 Dilemma 5; medium.com/@harsh_77214/beyond-naive-rag-comparing-basic-sentence-window-and-auto-merging-retrieval-....
MSC-05 — EmbedReturnDecouple (the core flip)
  • Trigger: Any time chunk-size tuning oscillates between precision and context.
  • Action: Set embed-unit small, return-unit large; never let them be the same object once the sweep is non-flat. Verify by inspecting what text the retriever actually sends to the synthesizer (must be the large unit).
  • Output: Pareto improvement on precision×context that no single size achieves.
  • Evidence: [[llamaindex]] Principle 2 (Node relationships); R3 Dilemma 1 resolution.
MSC-06 — MetadataBudgetGuard
  • Trigger: Small embed-units (≤256 tokens) with rich metadata propagated into payload.
  • Action: Ensure metadata never occupies >50% of the embed-unit token budget; strip/shorten metadata before shrinking chunks (GitHub #12200, #13792).
  • Output: Embed-units carry signal, not mostly boilerplate metadata.
  • Evidence: github.com/run-llama/llama_index/issues/12200, #13792; [[llamaindex]] A7.
MSC-07 — MeasureOrRevert
  • Trigger: Multi-scale config wired but lift unverified.
  • Action: Re-run the Gate-0 eval set; require a measurable faithfulness lift over the best single chunk size; if none, revert to the pinned single size.
  • Output: Evidence that the added complexity earns its keep — or its removal.
  • Evidence: [[llamaindex]] Stage 2 (eval loop gates every change).

5. 困境决策案例 (Dilemma Cases)

Show full SKILL.md (877 more words)Show less
5.1 — Sentence-Window vs Auto-Merging (which "embed-small-return-large"?)

困境: Both patterns implement the same core flip. They are not interchangeable — choosing wrong wastes setup cost and underperforms.

约束:

  • Sentence-Window expands horizontally — N adjacent sentences around the matched one.
  • Auto-Merging expands vertically — returns the parent when ≥ threshold child chunks match.
  • Sentence-Window has lower setup cost (one parser + one postprocessor).
  • Auto-Merging needs a docstore holding the full node hierarchy.

决策步骤:

  1. Documents have clear hierarchical structure (sections, headings, ToC) → Auto-Merging.
  2. Documents are flat narrative prose → Sentence-Window.
  3. Queries are bursty multi-chunk ("this entire section is relevant") → Auto-Merging escalates correctly.
  4. Queries are point-fact with surrounding context needed → Sentence-Window.
  5. Structure unknown → start Sentence-Window (lower setup cost), escalate if multi-chunk queries underperform.

结果: Both consistently beat naive top-k on faithfulness in published comparisons. Auto-Merging is more principled for structured docs; Sentence-Window is more robust for unstructured prose. The decision is driven by document structure, not theoretical elegance. (Source: [[llamaindex]] R3 Dilemma 5; developers.llamaindex.ai/.../auto_merging_retriever/; medium.com/@harsh_77214/beyond-naive-rag-comparing-basic-sentence-window-and-auto-merging-retrieval-...)

可提取的操作: MSC-02, MSC-03, MSC-04.

5.2 — Chunk-size sweep: the 1024 optimum, and when it does not converge

困境: At chunk_size=256 embeddings are precise but the LLM gets fragments; at chunk_size=2048 context is rich but the embedding becomes a "topic average" and recall on specific queries drops. Where to set chunk_size — and what to do when no single value wins?

约束:

  • Cannot test in production; need a deterministic offline answer.
  • Embedding model has a fixed input window (e.g. 512 tokens for many BGE variants — over-chunking is a hard error).
  • Synthesis-side token budget caps how many chunks fit downstream.
  • Metadata propagated into payload makes very small chunks "all metadata" (issues #12200, #13792).

决策步骤:

  1. Generate ~20 eval QA pairs.
  2. Sweep chunk_size ∈ {128,256,512,1024,2048}, overlap 10–20%.
  3. Build a VectorStoreIndex per config; record faithfulness + relevancy + latency.
  4. Single winner → pin it.
  5. Non-flat frontier → do not compromise; switch to embed-small/return-large via Sentence-Window or Auto-Merging.

结果: LlamaIndex's own published evaluation on Uber's 10-K found faithfulness peaked at chunk_size 1024 and relevancy maxed at 1024, with only mild latency growth — so 1024 became the framework default for prose (code lands at 80–160 tokens). But on corpora where the curve does not converge, the multi-scale decoupling pattern wins; never average two bad chunk sizes into one mediocre one. (Source: [[llamaindex]] R3 Dilemma 1; llamaindex.ai/blog/evaluating-the-ideal-chunk-size-for-a-rag-system-using-llamaindex-6207e5d3fec5; statsig.com/perspectives/llamaindex-rag-retrieval)

可提取的操作: MSC-01, MSC-05, MSC-06.


6. 反模式与边界 (Anti-patterns & Boundaries)

Anti-patterns
#Anti-patternCorrect move
A1Bump chunk_size (e.g. → 4096) when answers feel incompleteDecouple embed-scope from return-scope (MSC-02/03); do not enlarge the embed-unit
A2Use SentenceWindowNodeParser without MetadataReplacementPostProcessorAlways pair them — else the lone sentence, not the window, reaches the LLM
A3Index parent nodes in the vector store for auto-mergingIndex leaf nodes; keep parents in the docstore for merge-on-retrieval
A4Pick a single compromise chunk size on a non-flat frontierRefuse the compromise; switch to multi-scale (MSC-05)
A5Reach for multi-scale chunking on a short/static corpusPrompt-stuff with caching; multi-scale is over-engineering (B1)
A6Ship multi-scale config without re-running the eval setMSC-07: measure the lift or revert
A7Shrink the embed-unit while metadata still dominates the payloadMSC-06: budget metadata <50% before shrinking
Boundaries — when not to use multi-scale chunking
  • B1 — Short / static corpus (<100k tokens): prompt-stuff with caching; the chunk paradox does not arise. (Mirrors [[llamaindex]] B1.)
  • B2 — Sweep already converged: a single chunk size won both metrics → pin it and stop; multi-scale adds complexity with no payoff.
  • B3 — Bottleneck is elsewhere: if faithfulness is limited by the embedding model, the reranker, or the synthesizer (lost-in-the-middle), fix that first — multi-scale chunking only resolves the precision-vs-context axis.
  • B4 — Hard real-time (<100ms) retrieval: auto-merging's docstore lookups and window expansion add latency; a raw vector store may be the right tool.
PR-review smells (instant red flags)
  • SentenceSplitter(chunk_size=4096) introduced as a fix for "incomplete answers" → A1.
  • SentenceWindowNodeParser present but no MetadataReplacementPostProcessor in the query engine → A2.
  • AutoMergingRetriever over an index built from parent nodes (no leaf docstore) → A3.
  • Multi-scale parser added with no eval-set delta in the PR description → A6.

7. 跨框架对照 (Cross-framework Mapping)

The "embed small, return large" pattern is framework-agnostic; the primitives differ.

ConceptLlamaIndexLangChainNotes
Horizontal (sentence-window)SentenceWindowNodeParser + MetadataReplacementPostProcessor(no direct equivalent; emulate with custom retriever returning neighbor windows)LlamaIndex's is the cleanest first-class implementation
Vertical (parent-child / auto-merging)HierarchicalNodeParser + AutoMergingRetrieverParentDocumentRetriever (child splitter + parent splitter + docstore)Same idea: embed children, return parents
Small-embed unit storeVectorStoreIndex over leaf nodeschild vectorstoreboth index the small unit
Large-return unit storedocstore (nodes with PARENT/CHILD relationships)InMemoryStore / byte-store for parent docsthe return-unit lives outside the vector index

Mapping rule: LlamaIndex AutoMergingRetriever/HierarchicalNodeParser ≈ LangChain ParentDocumentRetriever. LlamaIndex additionally offers the horizontal SentenceWindowNodeParser, which LangChain has no first-class analogue for. For a coder agent already inside the LlamaIndex stack, prefer the native parsers; the [[llamaindex]] base skill governs the surrounding pipeline (ingestion, eval loop, reranking, routing).

This overlay does not replace [[llamaindex]] — it deepens the single DecoupleChunkScope knob into a full recipe. For everything around it (baseline, eval, hybrid, rerank, routing, production hardening), defer to the base skill.


References

  • references/R1-source-evidence.md — citations and provenance for every claim above.
  • intermediate/operation_candidates.json — machine-readable MSC operation list.
  • Base skill: [[llamaindex]] (SKILL.md + references/R3-dilemma-cases.md Dilemmas 1 & 5).
Primary sources (cited inline)
  • llamaindex.ai/blog/evaluating-the-ideal-chunk-size-for-a-rag-system-using-llamaindex-6207e5d3fec5 (Uber 10-K, 1024 optimum)
  • developers.llamaindex.ai/python/framework/integrations/retrievers/auto_merging_retriever/
  • developers.llamaindex.ai SentenceWindowNodeParser / MetadataReplacementPostProcessor / HierarchicalNodeParser docs
  • developers.llamaindex.ai/python/framework/optimizing/rag_failure_mode_checklist/ (failures #2, #6)
  • medium.com/@harsh_77214/beyond-naive-rag-comparing-basic-sentence-window-and-auto-merging-retrieval-with-llamaindex-f778173bed98
  • statsig.com/perspectives/llamaindex-rag-retrieval (code chunk size 80–160)
  • github.com/run-llama/llama_index/issues/12200, #13792 (metadata-dominates-chunk)
  • LangChain ParentDocumentRetriever docs (python.langchain.com/docs/how_to/parent_document_retriever/)

© agentsope, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (references) in skills/agentsop-multiscale-chunking of agentsope/SkillAlchemy.

  • SKILL.md
  • README.md
  • intermediate/operation_candidates.json
  • references/R1-source-evidence.md

Open the folder on GitHubat commit 6ea799f

Compare with similar skills

Agentsop Multiscale Chunking next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Agentsop Multiscale Chunking compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Agentsop Multiscale Chunking this skillagentsope/SkillAlchemy459—~4.9kAutomated safety check: PassMIT
Sentence Transformers EmbeddingsOrchestra-Research/AI-Research-SKILLs13k3 repos~1.6kAutomated safety check: PassMIT
FAISS Similarity SearchOrchestra-Research/AI-Research-SKILLs13k7 repos~1.3kAutomated safety check: PassMIT
RAG Skillsllama-farm/llamafarm8361 repos~1.3kAutomated safety check: PassApache-2.0
Chroma Vector DatabaseOrchestra-Research/AI-Research-SKILLs13k8 repos~2.3kAutomated safety check: PassMIT
Ms Agent Framework RAGshuyu-labs/WebCode278—~1.1kAutomated safety check: PassCustom licence

Similar skills

  • Sentence Transformers Embeddings

    Orchestra-Research/AI-Research-SKILLs

    Generates text embeddings locally with the sentence-transformers library for RAG, semantic search, clustering and similarity, with model picks for general, multilingual and legal text.

    13k GitHub starsUsed in 3 repos~1.6k tokens
    AI & LLM EngineeringAuto-check passed
  • FAISS Similarity Search

    Orchestra-Research/AI-Research-SKILLs

    Sets up FAISS for fast nearest-neighbor search over large collections of dense vectors, choosing between Flat, IVF, HNSW and product quantization indexes.

    13k GitHub starsUsed in 7 repos~1.3k tokens
    DatabasesAuto-check passed
  • RAG Skills

    llama-farm/llamafarm

    RAG-specific best practices for LlamaIndex, ChromaDB, and Celery workers.

    836 GitHub starsUsed in 1 repo~1.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Chroma Vector Database

    Orchestra-Research/AI-Research-SKILLs

    Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.

    13k GitHub starsUsed in 8 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Ms Agent Framework RAG

    shuyu-labs/WebCode

    Comprehensive guide for building Agentic RAG systems using Microsoft Agent Framework in C.

    278 GitHub stars~1.1k tokensUpdated 3 mo ago
    AI & LLM EngineeringAuto-check passed
  • Pgvector Semantic Search

    timescale/pg-aiguide

    A skill your agent uses for setting up vector similarity search with pgvector for AI/ML embeddings, RAG applications, or semantic search.

    1.9k GitHub starsUsed in 1 repo~3.8k tokens
    AI & LLM EngineeringAuto-check passed

More from agentsope/SkillAlchemy

All 45 skills in this repo
  • Agentsop Aider

    agentsope/SkillAlchemy

    SOP for terminal-based, git-native AI pair programming with Aider (git work-tree + tree-sitter repo-map + edit-format + human-in-loop REPL).

    459 GitHub stars~3.5k tokensUpdated 1 mo ago
    Auto-check passed
  • Agentsop Context Scope Discipline

    agentsope/SkillAlchemy

    Coder-agent working-file budget discipline: keep the editable working set (files you /add into writable context) under ~25k tokens, separate "read" from "edit", delegate breadth to a read-only…

    459 GitHub stars~3k tokensUpdated 1 mo ago
    Auto-check passed
  • Agentsop Cost Tiered Models

    agentsope/SkillAlchemy

    Split a multi-call LM workflow by cognitive load, not by accuracy: let one strong model make the few reasoning decisions and a cheap model do the many mechanical executions (Aider architect+editor…

    459 GitHub stars~3k tokensUpdated 1 mo ago
    Auto-check passed
  • Agentsop Crewai

    agentsope/SkillAlchemy

    SOP for building multi-agent systems with CrewAI — role-based collaboration, sequential/hierarchical processes, Flows, memory, delegation.

    459 GitHub stars~4.8k tokensUpdated 1 mo ago
    Auto-check passed
  • Agentsop Dify

    agentsope/SkillAlchemy

    SOP for building LLM applications on Dify — visual workflow + chatflow + agent + RAG knowledge base + plugin marketplace + observability, self-hostable.

    459 GitHub stars~5.4k tokensUpdated 1 mo ago
    Auto-check: notes
  • Agentsop Observability Setup

    agentsope/SkillAlchemy

    Enhancement-overlay skill — the DECISION + WIRING layer for LM observability that the single-backend skills [[langsmith]], [[phoenix]], [[mlflow]] do NOT cover.

    459 GitHub stars~4.4k tokensUpdated 1 mo ago
    Auto-check passed

Works with

Questions about Agentsop Multiscale Chunking

What does Agentsop Multiscale Chunking do?

Designs multiscale chunking for RAG by embedding small units for retrieval precision and returning larger context for synthesis. Agentsop Multiscale Chunking is an agent skill from agentsope/SkillAlchemy. Designs multiscale chunking for RAG by embedding small units for retrieval precision and returning larger context for synthesis.

When should I use Agentsop Multiscale Chunking?

Agentsop Multiscale Chunking fits situations like: fixed-size chunks either lose surrounding context; dilute relevance in long documents; A single chunk scale already meets retrieval and generation needs.

How do I install Agentsop Multiscale Chunking in Claude Code?

Run `npx skills add agentsope/SkillAlchemy --skill agentsop-multiscale-chunking -a claude-code`. Or copy the skill folder (skills/agentsop-multiscale-chunking in agentsope/SkillAlchemy) into .claude/skills/agentsop-multiscale-chunking in your project. Claude Code loads it when a task matches its description.

How do I install Agentsop Multiscale Chunking in Codex?

Run `npx skills add agentsope/SkillAlchemy --skill agentsop-multiscale-chunking -a codex`. Or copy the skill folder (skills/agentsop-multiscale-chunking in agentsope/SkillAlchemy) into .agents/skills/agentsop-multiscale-chunking in your project. Codex loads it when a task matches its description.

Can I use Agentsop Multiscale Chunking in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add agentsope/SkillAlchemy --skill agentsop-multiscale-chunking -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/agentsop-multiscale-chunking, .gemini/skills/agentsop-multiscale-chunking, .github/skills/agentsop-multiscale-chunking and .opencode/skills/agentsop-multiscale-chunking in your project.

What does Agentsop Multiscale Chunking need to run?

SKILL.md names no scripts, command-line tools or credentials: Agentsop Multiscale Chunking is instructions for the agent only. Our summary lists: Python 3.

Does Agentsop Multiscale Chunking access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Agentsop Multiscale Chunking safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Agentsop Multiscale Chunking use?

Agentsop Multiscale Chunking is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Agentsop Multiscale Chunking use?

About 4.9k tokens (SKILL.md is roughly 20k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.9k tokens, read only when the agent opens those files.

What are the alternatives to Agentsop Multiscale Chunking?

Skills that share tags, products or a category with Agentsop Multiscale Chunking: Sentence Transformers Embeddings (Orchestra-Research/AI-Research-SKILLs, 13k stars), FAISS Similarity Search (Orchestra-Research/AI-Research-SKILLs, 13k stars), RAG Skills (llama-farm/llamafarm, 836 stars) and Chroma Vector Database (Orchestra-Research/AI-Research-SKILLs, 13k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Agentsop Multiscale Chunking?

agentsope (a GitHub user) maintains it in agentsope/SkillAlchemy, which has 459 GitHub stars. The repository holds 45 skills in this directory. The repository was last updated on September 2, 2026.

Source: agentsope/SkillAlchemy on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.