Agent skill

Agentsop Llamaindex

by agentsope in agentsope/SkillAlchemy

Operating-system distillation of LlamaIndex — the leading RAG / document-agent framework.

MITAuto-check passedAI & LLM Engineering

Install Agentsop Llamaindex

skills CLI
$ npx skills add agentsope/SkillAlchemy --skill agentsop-llamaindex -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install agentsope/SkillAlchemy agentsop-llamaindex --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/agentsope/SkillAlchemy.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/agentsop-llamaindex .claude/skills/agentsop-llamaindex && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
agentsop-llamaindex
GitHub stars
466
Token cost
~6.4k tokens
SKILL.md length
2,558 words
Files
8 (incl. references)
Skills in repo
46
Repo updated
First seen
Licence
MIT

At a glance

Operating-system distillation of LlamaIndex — the leading RAG / document-agent framework.

  • Works in 7 steps: Frame the problem → Baseline (cheap, fast, observable) → Build the eval loop before optimizing… → …
  • Tasks that involve Building AI agents
  • SKILL.md covers 何时激活 (Activation Rules), 核心心智模型 (Core Mental Model), SOP 工作流 (Agentic Protocol) and 操作模型 (Operation Models), plus 1 more section
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Agentsop Llamaindex is an agent skill from agentsope/SkillAlchemy. Operating-system distillation of LlamaIndex — the leading RAG / document-agent framework. Activate when the calling agent must build, debug, harden, or evaluate a Retrieval-Augmented Generation pipeline over unstructured/private data, decide between RAG primitives (Index types, retrievers, query engines, routers, agents), or pick LlamaIndex vs LangChain / Haystack / raw vector store for a coding task. Encodes the 5-layer mental model (Documents → Nodes → Indices → Retrievers → Query Engines / Response…

Its SKILL.md is about 6.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 9 other files, including reference files (for example `README.md`, `intermediate/operation_candidates.json` and `references/R1-architecture.md`).

It sits in AI & LLM Engineering, covering Building AI agents, Retrieval-augmented generation and Operations and SOPs. It works with LlamaIndex, LangChain and GitHub. The repository describes itself as: From thought to skill. From signal to structure. The licence is MIT.

When your agent uses it

  • Tasks that involve Building AI agents
  • Tasks that involve Retrieval-augmented generation
  • Tasks that involve Operations and SOPs

Example prompts

  • “/agentsop-llamaindex”

Requirements

  • Python 3

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Frame the problem
  2. Baseline (cheap, fast, observable)
  3. Build the eval loop before optimizing anything
  4. Optimize in LlamaIndex's recommended order
  5. Compose for query heterogeneity
  6. Production hardening
  7. Escalate to Workflows / Agents (only when justified)

What it can do on your machine

Read from SKILL.md and the folder at commit d0f0355. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Agentsop Llamaindex loads about 6.4k tokens when it runs, and up to ~19k if it reads all its reference files. Until then it costs about 196 tokens; SKILL.md has 2,558 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~196
When it runs · the whole SKILL.md, loaded when a task matches
~6.4k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~19k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from agentsope/SkillAlchemy at commit d0f0355, republished under its MIT licence (© agentsope). 2,558 words, ~6,448 tokens.

Download SKILL.mdSave it as .claude/skills/agentsop-llamaindex/SKILL.md (or your agent's skills folder). This skill also uses 7 other files; get the full folder from GitHub.
name
agentsop-llamaindex
description
Operating-system distillation of LlamaIndex — the leading RAG / document-agent framework. Activate when the calling agent must build, debug, harden, or evaluate a Retrieval-Augmented Generation pipeline over unstructured/private data, decide between RAG primitives (Index types, retrievers, query engines, routers, agents), or pick LlamaIndex vs LangChain / Haystack / raw vector store for a coding task. Encodes the 5-layer mental model (Documents → Nodes → Indices → Retrievers → Query Engines / Response Synthesizers), the canonical RAG bootstrap SOP from baseline `VectorStoreIndex` through hybrid + reranker + eval-loop hardening, the official 13-failure-mode checklist, and 5 dilemma cases distilled from docs, GitHub issues, and 2025 production post-mortems.
version
0.1.0

LlamaIndex · SOP

Third-person analytical view of how LlamaIndex thinks about turning private documents into a grounded answering system. The skill is for an LLM agent that writes / reviews / debugs RAG code — not for an end user reading docs.


何时激活 (Activation Rules)

Activate this skill when any of the following holds:

  1. The user's request involves building, modifying, or debugging a RAG pipeline (retrieval over private/unstructured data + LLM synthesis).
  2. The user mentions LlamaIndex (from llama_index...), LlamaParse, LlamaCloud, or a LlamaIndex-style primitive (VectorStoreIndex, SummaryIndex, IngestionPipeline, QueryEngine, SubQuestionQueryEngine, RouterQueryEngine, Settings, Workflows).
  3. The user is comparing RAG frameworks (LlamaIndex vs LangChain vs Haystack vs raw vector store).
  4. The user is choosing between stuffing context, RAG, or an agent for a knowledge task.
  5. The user is debugging retrieval quality (hallucinations, wrong chunks, stale data, embedding drift) — even if the codebase predates LlamaIndex, the failure-mode taxonomy applies.
  6. The user is evaluating a RAG system (faithfulness, relevancy, MRR, hit-rate).

Do not activate when:

  • The task is pure agent orchestration with no retrieval (use LangGraph/CrewAI skill instead).
  • The corpus is tiny (<100k tokens, static) and prompt-stuffing is the correct answer.
  • The data is pure SQL/tabular with no unstructured component.

核心心智模型 (Core Mental Model)

LlamaIndex's design rests on three principles that distinguish it from "vector DB SDK + custom glue":

Principle 1 — The Index is a noun, not a verb

In LangChain, "indexing" is something you do to a vector store. In LlamaIndex, an Index is a first-class typed object with its own retrieval semantics. Picking the right Index is half the architecture decision.

The 5-layer pipeline:

Documents → Nodes → Index → Retriever → Query Engine → Response
   ↓         ↓        ↓         ↓             ↓
parsing   chunking  storage   filters    synthesis
metadata  graph     primitive  rerank    (refine/tree_sum/compact)

Each layer has a distinct failure mode and a distinct optimization knob. See references/R1-architecture.md for the layer-failure-knob mapping.

Principle 2 — A Node is a graph node, not a chunk

A Node carries: text, metadata, embedding, relationships (PREV/NEXT/PARENT/CHILD links), and lifecycle ids. The relationships field is what enables Hierarchical, Auto-Merging, and Sentence-Window retrieval. The mental flip: don't think "split into chunks", think "build a chunk-graph".

Principle 3 — Indices are not interchangeable
IndexPick when
VectorStoreIndexDefault; semantic Q&A over chunks; ~90% of RAG cases
SummaryIndex"Summarize this whole doc" — small, fan-out synthesis
TreeIndexHierarchical content with progressive zoom-in
KeywordTableIndexKeyword-heavy queries, no embeddings budget
PropertyGraphIndexMulti-hop reasoning over entities
DocumentSummaryIndexMixed corpora needing document-level routing first

A RouterQueryEngine over multiple per-task indices is often the correct top-level shape, not a single monolithic VectorStoreIndex.

The 2025 shift

LlamaIndex now positions as "the leading document agent and OCR platform" (README). LlamaParse v2 + Workflows 1.0 (June 2025) + LlamaCloud mark a strategic move from "RAG framework" to "platform between messy documents and document-grounded agents". For a coder agent: assume Workflows for any new agentic code (QueryPipeline is deprecated).


SOP 工作流 (Agentic Protocol)

The protocol every RAG implementation must walk through. Each stage gates on the next.

Stage 0 — Frame the problem

Before code, answer:

  1. Is the corpus unstructured + non-trivial size (>100k tokens) + growing? If not → see R4 boundaries; LlamaIndex may be the wrong tool.
  2. Is retrieval quality the bottleneck (not orchestration)? If orchestration dominates → LangGraph leads, LlamaIndex becomes a retrieval tool inside it.
  3. What is the query distribution? (lookup-only / summary / compare-contrast / mixed). This decides whether a single Index or a Router is needed.
Stage 1 — Baseline (cheap, fast, observable)
python
from llama_index.core import VectorStoreIndex, SimpleDirectoryReader, Settings
from llama_index.core.node_parser import SentenceSplitter

Settings.llm        = OpenAI(model="gpt-4o-mini")
Settings.embed_model = OpenAIEmbedding(model="text-embedding-3-small")
Settings.node_parser = SentenceSplitter(chunk_size=1024, chunk_overlap=20)

docs  = SimpleDirectoryReader("./data").load_data()
index = VectorStoreIndex.from_documents(docs)
qe    = index.as_query_engine(similarity_top_k=4)

Pin Settings once at app boot, never inline. This eliminates the entire embedding-mismatch failure class (failure #4).

Stage 2 — Build the eval loop before optimizing anything
python
from llama_index.core.evaluation import (
    DatasetGenerator, FaithfulnessEvaluator,
    RelevancyEvaluator, RetrieverEvaluator,
)
qa = DatasetGenerator.from_documents(docs).generate_dataset_from_nodes(num=50)

Track {MRR, hit-rate, faithfulness, relevancy, p95 latency}. Every subsequent change must be gated on these numbers.

Most RAG failures in production trace to weak retrieval or sloppy ingestion — not the LLM. The eval loop is what surfaces them.

From the official basic_strategies guide:

  1. Prompt engineering (cheapest)
  2. Embedding model (pick from MTEB; full re-embed if you change it)
  3. Chunk size sweep ({256, 512, 1024, 2048}; default 1024 for prose, 80-160 for code)
  4. Hybrid search (BM25 + dense) — only if traffic contains lexical-identity queries
  5. Metadata filters — for multi-tenant / multi-collection corpora
  6. Document/chunk decoupling — HierarchicalNodeParser+AutoMergingRetriever or SentenceWindowNodeParser
  7. Reranking (Cohere / SentenceTransformer / ColBERT) — widen top_k to 20-50, rerank to 3-5

Note the order: prompts first, reranking last. Reranking is high-impact but expensive — exhaust cheap knobs first.

Stage 4 — Compose for query heterogeneity
Query shapeRight primitive
"Summarize doc X"SummaryIndex per doc, routed
"Find the clause about X"VectorStoreIndex + metadata filters
"Compare X and Y across docs"SubQuestionQueryEngine
"What entities relate to X?"PropertyGraphIndex
MixedRouterQueryEngine over per-task engines
Stage 5 — Production hardening

Apply the failure-mode checklist (R4). Top 5 non-negotiables:

  • IngestionPipeline with docstore + UPSERTS_AND_DELETE for any live corpus.
  • Settings.embed_model pinned at boot; embedding model name in index metadata.
  • tree_summarize synthesizer when packing many chunks (mitigates lost-in-the-middle).
  • Tracing/observability captures query + retrieved_nodes + scores + index_id + LLM prompt for every failure.
  • Indices versioned as deployment artifacts; ingestion completes before traffic routing.
Stage 6 — Escalate to Workflows / Agents (only when justified)

Escalate when at least one of:

  • A retrieval loop is needed ("retrieve → check → re-query").
  • Tool calls beyond retrieval (calculator, web, code-exec).
  • State surviving across query turns.
  • Multiple specialized retrievers chosen at runtime.

Use Workflows 1.0 (event-driven), not deprecated QueryPipeline. Wrap query engines as QueryEngineTools and tune the description= carefully — it is the only signal the router/agent reads.


操作模型 (Operation Models)

Each operation: Trigger / Action / Output / Evidence.

OP-01 BaselineVectorIndex
  • Trigger: First-pass RAG over a new corpus; retrieval-quality baseline unknown.
  • Action: VectorStoreIndex.from_documents() with SentenceSplitter(1024, 20), top_k=4, default synthesizer. Ship to eval bench before tuning.
  • Output: Working RAG endpoint + baseline {MRR, hit-rate, faithfulness, relevancy, p95}.
  • Evidence: developers.llamaindex.ai/python/framework/optimizing/basic_strategies/basic_strategies/
OP-02 TuneChunkSize
  • Trigger: Faithfulness below target OR retrieved chunks visibly truncated/incomplete.
  • Action: Sweep chunk_size ∈ {256, 512, 1024, 2048} with overlap at ~10-20%; re-evaluate faithfulness + relevancy + latency. Default land: 1024 for prose, 80-160 for code.
  • Output: Optimal chunk_size pinned + embedding model version locked in index metadata.
  • Evidence: llamaindex.ai/blog/evaluating-the-ideal-chunk-size-for-a-rag-system-using-llamaindex-6207e5d3fec5 (faithfulness peaked at 1024 in LlamaIndex's own eval on Uber 10-K).
OP-03 AddReranker
  • Trigger: Top-1 wrong but relevant docs appear in top-k (failure #1 / #10).
  • Action: Add CohereRerank or SentenceTransformerRerank as a NodePostprocessor; widen retrieval top_k to 20-50, narrow to top_n=3-5 after rerank.
  • Output: Faithfulness lift typically 5-15pp on noisy corpora; lower context-window pressure.
  • Evidence: developers.llamaindex.ai/python/framework/optimizing/rag_failure_mode_checklist/ (#1, #10).
OP-04 AddHybridBM25
  • Trigger: Traffic contains exact identifiers, error codes, SKUs, code symbols, rare jargon — pure dense silently misses them.
  • Action: QueryFusionRetriever([vector_retriever, BM25Retriever]) or vendor hybrid (Qdrant/Milvus alpha). Tune alpha per query type, not globally.
  • Output: Recall lift on lexical-identity queries with no degradation on semantic queries.
  • Evidence: llamaindex.ai/blog/llamaindex-enhancing-retrieval-performance-with-alpha-tuning-in-hybrid-search-in-rag-135d0c9b8a00; BM25Retriever docs.
OP-05 DecoupleChunkScope
  • Trigger: Chunk-size sweep produces no single winner (small wins precision, large wins context).
  • Action: HierarchicalNodeParser + AutoMergingRetriever (for structured docs) or SentenceWindowNodeParser + MetadataReplacementPostProcessor (for flat prose). Embed small, return large.
  • Output: Precision-recall pareto improvement; LLM gets surrounding context that small chunks alone lost.
  • Evidence: AutoMergingRetriever / Hierarchical / SentenceWindow docs on developers.llamaindex.ai.
OP-06 RouteByQueryType
  • Trigger: Corpus serves heterogeneous tasks (summary / lookup / compare) from one entry point.
  • Action: Build per-task QueryEngines (SummaryIndex for digest, VectorStoreIndex for lookup, SubQuestionQueryEngine for compare) + a RouterQueryEngine with LLM or Pydantic selector. Carefully author each QueryEngineTool.description.
  • Output: Each query lands on the structurally-correct retrieval primitive; latency stays bounded.
  • Evidence: DeepLearning.AI Building Agentic RAG with LlamaIndex; router docs.
OP-07 DecomposeMultiHop
  • Trigger: Compare/contrast queries; queries needing facts from >1 document; "what changed between X and Y?".
  • Action: SubQuestionQueryEngine decomposes query → dispatches sub-questions to sub-engines → synthesizes.
  • Output: Multi-hop answers a single retrieval cannot assemble.
  • Evidence: developers.llamaindex.ai sub-question query engine docs.
OP-08 IngestionWithDocstore
  • Trigger: Documents will update/delete over time (any production system).
  • Action: IngestionPipeline(transformations=..., docstore=..., vector_store=..., docstore_strategy=UPSERTS_AND_DELETE). Run on a schedule, not manually.
  • Output: Idempotent re-ingestion; no duplicate vectors; deletes propagate.
  • Evidence: developers.llamaindex.ai/python/framework/module_guides/loading/ingestion_pipeline/; failure #3.
OP-09 MetadataFilters
  • Trigger: Multi-tenant corpus; cross-contamination between sub-collections; access control needed.
  • Action: Inject structured metadata at ingestion (tenant, doc_type, date); apply MetadataFilters at query time OR enable auto-retrieval to let an LLM emit filters.
  • Output: Hard isolation between tenants; targeted retrieval without expensive rerank.
  • Evidence: Failure #7; basic_strategies metadata filters section.
OP-10 EvalLoop
  • Trigger: Any non-trivial RAG, pre-deploy AND continuously in production.
  • Action: DatasetGenerator → labeled QA pairs; run FaithfulnessEvaluator + RelevancyEvaluator + RetrieverEvaluator(["mrr","hit_rate"]). Gate every change.
  • Output: Quantitative regression test for every chunking / embedding / retriever / prompt change.
  • Evidence: developers.llamaindex.ai/python/framework-api-reference/evaluation/; cookbook.openai.com/examples/evaluation/evaluate_rag_with_llamaindex.
OP-11 LockGlobalSettings
  • Trigger: Multiple modules each instantiate LLM/embed independently — drift risk.
  • Action: Set Settings.llm and Settings.embed_model once in app bootstrap. Forbid inline overrides in PR review.
  • Output: Eliminates failure #4 (config drift) and #5 (embedding mismatch).
  • Evidence: docs.llamaindex.ai/en/stable/module_guides/supporting_modules/service_context_migration/.
OP-12 AgenticWorkflow
  • Trigger: Need loops, tool calls beyond retrieval, multi-step reasoning, or state across turns.
  • Action: Build a Workflows 1.0 event-driven workflow OR a FunctionAgent/ReActAgent with QueryEngineTools. Do NOT use the deprecated QueryPipeline.
  • Output: Cycle-capable agentic system with retrieval as one tool among many.
  • Evidence: llamaindex.ai/blog/announcing-workflows-1-0-a-lightweight-framework-for-agentic-systems.

困境决策案例 (Dilemma Cases)

(Full text in references/R3-dilemma-cases.md. Summarized here.)

Dilemma 1 — Chunk size: precision vs context

困境: Small chunks → precise embeddings, fragmented context for the LLM. Large chunks → rich context, embeddings become "topic averages", recall on specific queries drops. Failure modes #2 and #6 are the two poles.

约束: Embedding model has a fixed input window; metadata is propagated into payload (so very small chunks become all-metadata — GitHub #12200, #13792); token budget caps how many chunks fit downstream.

决策步骤:

  1. Generate ~20 eval QA pairs.
  2. Sweep chunk_size ∈ {128, 256, 512, 1024, 2048} with overlap = 10-20%.
  3. Build a VectorStoreIndex per config; record faithfulness, relevancy, latency.
  4. If a single winner emerges → pin it.
  5. If the frontier is non-flat → do not compromise; switch to small-embed/large-return via Hierarchical+AutoMerging or SentenceWindow.

结果: LlamaIndex's own published study (Uber 10-K) peaked at 1024 on both faithfulness and relevancy → 1024 became the framework default for prose. For code: 80-160 tokens. When the eval doesn't converge, decoupling wins; never average two bad chunk_sizes.

可提取的操作: OP-02 TuneChunkSize, OP-05 DecoupleChunkScope. Anti-pattern A1.

Show full SKILL.md (959 more words)Show less
Dilemma 2 — Hybrid (BM25+dense) vs pure dense

困境: Adding hybrid doubles index footprint, requires per-query-type alpha tuning, complicates the pipeline. Worth it?

约束: Dense embeddings silently fail on identifiers, error strings, code, SKUs — they "destroy lexical identity by pooling token representations" (TianPan, 2026). BM25 scores against an inverted token index.

决策步骤:

  1. Build a query-type taxonomy from real traffic: semantic / lexical / mixed.
  2. If lexical share <5% → dense-only.
  3. 5-50% → add hybrid; tune alpha per query type.
  4. 50% (legal, code, logs) → invert: BM25-first, dense as reranker signal.

  5. Evaluate alpha at {0, 0.25, 0.5, 0.75, 1.0} on labeled subsets.

结果: Hybrid lifts the lexical slice without hurting the semantic slice — if alpha is tuned per type. A single global alpha often underperforms dense, which is why some teams wrongly conclude "hybrid didn't help".

可提取的操作: OP-04 AddHybridBM25. Decision is traffic-driven, not theoretical.

Dilemma 3 — Agent on top of RAG, RAG as tool, or just a Router?

困境: User adds compare/summary/lookup queries to a basic RAG. Three options:

  • A. RouterQueryEngine over per-task engines.
  • B. FunctionAgent/ReActAgent with engines as tools.
  • C. SubQuestionQueryEngine to decompose.

约束: Agents add ≥1 LLM round-trip per step (latency); introduce planning errors a router cannot make; harder to debug (failure #12); most queries aren't multi-hop in practice.

决策步骤:

  1. Measure: what fraction of queries actually need multi-step reasoning?
  2. <20% multi-step + heterogeneous-but-single-step → Router (A).
  3. Compositional/well-shaped queries ("compare X and Y") → SubQuestion (C).
  4. Tool calls beyond retrieval, or cycles, or state → Agent on Workflows (B).
  5. Whichever you pick: invest in QueryEngineTool.description — it's the only signal the router/agent sees.

结果: DeepLearning.AI's official course ladder is Router → Agent. Production guidance consistently warns against premature agentization. Workflows 1.0 (2025) signals: when you need agency, use the agentic primitive, don't fake it with DAG pipelines.

可提取的操作: OP-06 RouteByQueryType, OP-07 DecomposeMultiHop, OP-12 AgenticWorkflow. Anti-pattern A9.

Dilemma 4 — Long-context LLM (1M tokens) vs RAG

困境: Does a 1M-token context window eliminate the need for RAG?

约束 (from llamaindex.ai/blog/towards-long-context-rag): 1M tokens ~60s latency + $0.50-$20/query; 10M tokens still doesn't cover large corpora; "lost in the middle" degrades quality by ~30%.

决策步骤:

  1. Corpus >1M tokens → RAG mandatory.
  2. p50 latency budget <5s → cannot afford full-context stuffing.
  3. Per-query cost ceiling <$0.05 → same.
  4. Apply LlamaIndex's three long-context patterns: Small-to-Big, Intelligent Routing, Retrieval-Augmented KV Caching.

结果: Long context does not replace RAG; it changes what RAG looks like. The bottleneck shifts from "fitting context" to "feeding right context in the right position" — making rerank + position-aware synthesis (tree_summarize) more important, not less.

可提取的操作: For any corpus >500k tokens or latency <5s: keep RAG. Use long-context as synthesis-stage capacity.

Dilemma 5 — Sentence-Window vs Auto-Merging

困境: Both implement "embed small, return large". Not interchangeable.

决策步骤:

  1. Docs have clear hierarchy (sections/headings) → Auto-Merging.
  2. Docs are flat prose → Sentence-Window.
  3. Queries are bursty multi-chunk → Auto-Merging escalates correctly.
  4. Queries are point-fact with surrounding context → Sentence-Window.

结果: Both beat naive top-k on faithfulness. Match parser/retriever pair to document structure, not theoretical elegance. Always pair SentenceWindowNodeParser with MetadataReplacementPostProcessor.


反模式与边界 (Anti-patterns & Boundaries)

Top 10 anti-patterns (full list in references/R4-anti-patterns.md)
#Anti-patternCorrect move
A1Bump chunk_size when answers feel incompleteDecouple embed-scope from synthesis-scope (Hierarchical / SentenceWindow)
A2Swap embedding model without re-embedRebuild index; tag artifact with embed model name+version
A3No eval loop; debug by anecdoteStand up RetrieverEvaluator + FaithfulnessEvaluator + RelevancyEvaluator first
A4ServiceContext + manual config in every modulePin Settings.llm and Settings.embed_model once at boot
A5QueryPipeline DAG for agentic logicUse Workflows 1.0 (event-driven, supports cycles)
A6Naive top_k=N, no rerankerWiden top_k + add CohereRerank / SentenceTransformerRerank
A7Metadata not propagated to chunks; or metadata > 50% of chunk_sizeDesign metadata schema before ingestion; budget metadata tokens
A8Multi-modal RAG by base64-stuffing images into textUse LlamaParse + multi-modal retrieval primitives
A9Wrap retrieval in a custom agent when a Router sufficesDefault to RouterQueryEngine; escalate to Agent only with justification
A10Ingest once at deploy, never reconcileIngestionPipeline + docstore + UPSERTS_AND_DELETE
Boundaries — when not to use LlamaIndex
  • B1: Tiny static corpus (<100k tokens) → prompt-stuff with caching.
  • B2: Pure structured/tabular data → DuckDB/SQL/BI. (LlamaIndex only when NL2SQL+RAG hybrid.)
  • B3: Hard real-time / sub-100ms retrieval → raw vector store SDK, not a RAG framework.
  • B4: Complex multi-agent orchestration → LangGraph or CrewAI leads; embed LlamaIndex retrievers as tools.
  • B5: Highly specialized parsing requirements + team has engineering budget → custom stack (Unstructured.io + pgvector + custom retriever) gives more control.
PR-review smells (instant red flags)
  • from llama_index import ServiceContext → A4.
  • index.as_query_engine(similarity_top_k=20) without a rerank postprocessor → A6.
  • SentenceSplitter(chunk_size=4096) → likely A1.
  • Settings.embed_model = ... in >1 file → A4 drift.
  • IngestionPipeline(...) without docstore= → A10.
  • A Workflow with no events or loops → over-engineered; should be a QueryEngine.
  • An agent with a single retrieval tool → A9; should be a QueryEngine or RouterQueryEngine.

生态对照 (Ecosystem Context)

Decision rubric
Q1. Primarily extracting from messy documents (PDFs, slides, tables, scans)?
   YES → LlamaIndex (+ LlamaParse) leads.
Q2. Primary challenge is multi-step agentic orchestration with many non-retrieval tools?
   YES → LangGraph / CrewAI leads; use LlamaIndex retrievers as tools.
Q3. Corpus small (<100k tokens) and static?
   YES → No framework; prompt-stuff with caching.
Q4. Pure structured/tabular data?
   YES → SQL/DuckDB/BI. Use LlamaIndex only for hybrid NL2SQL+RAG.
DEFAULT → LlamaIndex remains lead; layer LangGraph only if agentic logic emerges.
Head-to-head highlights
VsLlamaIndex wins whenOther wins when
LangChainRetrieval quality and ingestion are the bottleneck; document-heavyOrchestration is complex; many non-retrieval tools
HaystackModern LLM-centric docs; multi-modal; broader index taxonomyYAML-configurable pipelines; classical IR feel
Raw vector storeNeed >2 of {SentenceSplitter, IngestionPipeline, Reranker, Eval, Synthesizer}Truly minimal RAG; team wants no framework
DSPyWant structured retrieval infrastructureWant automatic prompt optimization
LangGraph (for agents)Retrieval-heavy with light agency (Workflows ergonomic here)Many states, complex multi-agent state machines
CrewAI / AutoGen(different category)Multi-agent collaboration is the goal
The normative hybrid (2025-2026)

Most production teams converge on: LlamaIndex for retrieval & ingestion; LangGraph (or LlamaIndex Workflows) for orchestration; LangSmith / Phoenix for observability.


References

  • references/R1-architecture.md — 5-layer model deep dive, Index taxonomy, Settings/Workflows
  • references/R2-sop-workflow.md — full 8-stage RAG bootstrap protocol
  • references/R3-dilemma-cases.md — 5 dilemma cases in full
  • references/R4-anti-patterns.md — 13 official failure modes + 10 anti-patterns + boundaries
  • references/R5-ecosystem-context.md — comparison matrix, hybrid patterns
  • intermediate/operation_candidates.json — machine-readable operation list
Primary sources (cited inline above)
  • developers.llamaindex.ai/python/framework/ (architecture homepage)
  • developers.llamaindex.ai/python/framework/optimizing/basic_strategies/basic_strategies/
  • developers.llamaindex.ai/python/framework/optimizing/rag_failure_mode_checklist/ (official 13 failure modes)
  • developers.llamaindex.ai/python/framework/module_guides/indexing/index_guide/
  • developers.llamaindex.ai/python/framework/module_guides/loading/ingestion_pipeline/
  • llamaindex.ai/blog/evaluating-the-ideal-chunk-size-for-a-rag-system-using-llamaindex-6207e5d3fec5
  • llamaindex.ai/blog/llamaindex-enhancing-retrieval-performance-with-alpha-tuning-in-hybrid-search-in-rag-135d0c9b8a00
  • llamaindex.ai/blog/towards-long-context-rag
  • llamaindex.ai/blog/announcing-workflows-1-0-a-lightweight-framework-for-agentic-systems
  • docs.llamaindex.ai/en/stable/module_guides/supporting_modules/service_context_migration/
  • github.com/run-llama/llama_index (README, issues #12200, #13792, #6465)
  • cookbook.openai.com/examples/evaluation/evaluate_rag_with_llamaindex
  • learn.deeplearning.ai/courses/building-agentic-rag-with-llamaindex/
  • ibm.com/think/topics/llamaindex-vs-langchain
  • statsig.com/perspectives/llamaindex-rag-retrieval
  • tianpan.co/blog/2026-04-12-hybrid-search-production-bm25-dense-embeddings

© agentsope, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 7 other files (references) in skills/agentsop-llamaindex of agentsope/SkillAlchemy.

  • SKILL.md
  • README.md
  • intermediate/operation_candidates.json
  • references/R1-architecture.md
  • references/R2-sop-workflow.md
  • references/R3-dilemma-cases.md
  • references/R4-anti-patterns.md
  • references/R5-ecosystem-context.md

Open the folder on GitHubat commit d0f0355

Compare with similar skills

Agentsop Llamaindex next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Agentsop Llamaindex compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Agentsop Llamaindex this skillagentsope/SkillAlchemy466—~6.4kAutomated safety check: PassMIT
Neo4j Graphrag Skillneo4j-contrib/neo4j-skills114—~4.2kAutomated safety check: NotesMIT
LangchainOrchestra-Research/AI-Research-SKILLs13k2 repos~3.2kAutomated safety check: PassMIT
Langchain RAGlangchain-ai/langchain-skills1.3k—~3.9kAutomated safety check: PassMIT
Pinecone Vector DatabaseOrchestra-Research/AI-Research-SKILLs13k5 repos~2kAutomated safety check: PassMIT
Neo4j Document Import Skillneo4j-contrib/neo4j-skills114—~5.4kAutomated safety check: NotesMIT

Similar skills

  • Neo4j Graphrag Skill

    neo4j-contrib/neo4j-skills

    Build GraphRAG retrieval pipelines on Neo4j using the neo4j-graphrag Python package (v1.22.0+).

    114 GitHub stars~4.2k tokensUpdated 3 days ago
    Knowledge ManagementAuto-check: notes
  • Langchain

    Orchestra-Research/AI-Research-SKILLs

    Framework for building LLM-powered applications with agents, chains, and RAG.

    13k GitHub starsUsed in 2 repos~3.2k tokens
    AI & LLM EngineeringAuto-check passed
  • Langchain RAG

    langchain-ai/langchain-skills

    Official

    INVOKE THIS SKILL when building ANY retrieval-augmented generation (RAG) system.

    1.3k GitHub stars~3.9k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Pinecone Vector Database

    Orchestra-Research/AI-Research-SKILLs

    Shows how to use Pinecone, a managed vector database, for production RAG, semantic search and recommendations: indexes, upserts, queries, filters and namespaces.

    13k GitHub starsUsed in 5 repos~2k tokens
    DatabasesAuto-check passed
  • Neo4j Document Import Skill

    neo4j-contrib/neo4j-skills

    Ingests unstructured and semi-structured documents into Neo4j as a knowledge graph.

    114 GitHub stars~5.4k tokensUpdated 3 days ago
    Knowledge ManagementAuto-check: notes
  • FAISS Similarity Search

    Orchestra-Research/AI-Research-SKILLs

    Sets up FAISS for fast nearest-neighbor search over large collections of dense vectors, choosing between Flat, IVF, HNSW and product quantization indexes.

    13k GitHub starsUsed in 6 repos~1.3k tokens
    DatabasesAuto-check passed

More from agentsope/SkillAlchemy

All 46 skills in this repo
  • Agentsop Aider

    agentsope/SkillAlchemy

    SOP for terminal-based, git-native AI pair programming with Aider (git work-tree + tree-sitter repo-map + edit-format + human-in-loop REPL).

    466 GitHub stars~3.5k tokensUpdated yesterday
    Auto-check passed
  • Agentsop Context Scope Discipline

    agentsope/SkillAlchemy

    Coder-agent working-file budget discipline: keep the editable working set (files you /add into writable context) under ~25k tokens, separate "read" from "edit", delegate breadth to a read-only…

    466 GitHub stars~3k tokensUpdated yesterday
    Auto-check passed
  • Agentsop Cost Tiered Models

    agentsope/SkillAlchemy

    Split a multi-call LM workflow by cognitive load, not by accuracy: let one strong model make the few reasoning decisions and a cheap model do the many mechanical executions (Aider architect+editor…

    466 GitHub stars~3k tokensUpdated yesterday
    Auto-check passed
  • Agentsop Crewai

    agentsope/SkillAlchemy

    SOP for building multi-agent systems with CrewAI — role-based collaboration, sequential/hierarchical processes, Flows, memory, delegation.

    466 GitHub stars~4.8k tokensUpdated yesterday
    Auto-check passed
  • Agentsop Dify

    agentsope/SkillAlchemy

    SOP for building LLM applications on Dify — visual workflow + chatflow + agent + RAG knowledge base + plugin marketplace + observability, self-hostable.

    466 GitHub stars~5.4k tokensUpdated yesterday
    Auto-check: notes
  • Agentsop Multiscale Chunking

    agentsope/SkillAlchemy

    Designs multiscale chunking for RAG by embedding small units for retrieval precision and returning larger context for synthesis.

    466 GitHub stars~4.9k tokensUpdated yesterday
    Auto-check passed

Questions about Agentsop Llamaindex

What does Agentsop Llamaindex do?

Operating-system distillation of LlamaIndex — the leading RAG / document-agent framework. Agentsop Llamaindex is an agent skill from agentsope/SkillAlchemy. Operating-system distillation of LlamaIndex — the leading RAG / document-agent framework.

When should I use Agentsop Llamaindex?

Agentsop Llamaindex fits situations like: tasks that involve Building AI agents; tasks that involve Retrieval-augmented generation; tasks that involve Operations and SOPs.

How do I install Agentsop Llamaindex in Claude Code?

Run `npx skills add agentsope/SkillAlchemy --skill agentsop-llamaindex -a claude-code`. Or copy the skill folder (skills/agentsop-llamaindex in agentsope/SkillAlchemy) into .claude/skills/agentsop-llamaindex in your project. Claude Code loads it when a task matches its description.

How do I install Agentsop Llamaindex in Codex?

Run `npx skills add agentsope/SkillAlchemy --skill agentsop-llamaindex -a codex`. Or copy the skill folder (skills/agentsop-llamaindex in agentsope/SkillAlchemy) into .agents/skills/agentsop-llamaindex in your project. Codex loads it when a task matches its description.

Can I use Agentsop Llamaindex in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add agentsope/SkillAlchemy --skill agentsop-llamaindex -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/agentsop-llamaindex, .gemini/skills/agentsop-llamaindex, .github/skills/agentsop-llamaindex and .opencode/skills/agentsop-llamaindex in your project.

What does Agentsop Llamaindex need to run?

SKILL.md names no scripts, command-line tools or credentials: Agentsop Llamaindex is instructions for the agent only. Our summary lists: Python 3.

Does Agentsop Llamaindex access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Agentsop Llamaindex safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Agentsop Llamaindex use?

Agentsop Llamaindex is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Agentsop Llamaindex use?

About 6.4k tokens (SKILL.md is roughly 26k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 13k tokens, read only when the agent opens those files.

What are the alternatives to Agentsop Llamaindex?

Skills that share tags, products or a category with Agentsop Llamaindex: Neo4j Graphrag Skill (neo4j-contrib/neo4j-skills, 114 stars), Langchain (Orchestra-Research/AI-Research-SKILLs, 13k stars), Langchain RAG (langchain-ai/langchain-skills, 1.3k stars) and Pinecone Vector Database (Orchestra-Research/AI-Research-SKILLs, 13k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Agentsop Llamaindex?

agentsope (a GitHub user) maintains it in agentsope/SkillAlchemy, which has 466 GitHub stars. The repository holds 46 skills in this directory. The repository was last updated on October 9, 2026.

Source: agentsope/SkillAlchemy on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.