Agent skill

RAG Skills

by llama-farm in llama-farm/llamafarm

RAG-specific best practices for LlamaIndex, ChromaDB, and Celery workers.

Apache-2.0Auto-check passedAI & LLM Engineering

Install RAG Skills

skills CLI
$ npx skills add llama-farm/llamafarm --skill rag-skills -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install llama-farm/llamafarm rag-skills --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/llama-farm/llamafarm.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/rag-skills .claude/skills/rag-skills && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
rag-skills
GitHub stars
836
Token cost
~1.3k tokens
SKILL.md length
168 words
Files
5
Skills in repo
19
Repo updated
First seen
Licence
Apache-2.0

At a glance

RAG-specific best practices for LlamaIndex, ChromaDB, and Celery workers.

  • Works in 4 steps: LlamaIndex (Medium priority) → ChromaDB (High priority) → Celery (High priority) → …
  • Tasks that involve Retrieval-augmented generation
  • SKILL.md covers Component Overview, Directory Structure, Quick Reference and Core Patterns, plus 1 more section
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

RAG Skills is an agent skill from llama-farm/llamafarm. RAG-specific best practices for LlamaIndex, ChromaDB, and Celery workers. Covers ingestion, retrieval, embeddings, and performance.

Its SKILL.md is about 1.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files (for example `celery.md`, `chromadb.md` and `llamaindex.md`).

It sits in AI & LLM Engineering, covering Retrieval-augmented generation, Background jobs and Embeddings. It works with LlamaIndex, Chroma and Python. The repository describes itself as: Deploy any AI model, agent, database, RAG, and pipeline locally or remotely in minutes. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Retrieval-augmented generation
  • Tasks that involve Background jobs
  • Tasks that involve Embeddings

Example prompts

  • “/rag-skills”

Requirements

  • Python 3
  • Pre-approved tools (allowed-tools): Read, Grep, Glob

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. LlamaIndex (Medium priority)
  2. ChromaDB (High priority)
  3. Celery (High priority)
  4. Performance (Medium priority)

What it can do on your machine

Read from SKILL.md and the folder at commit 6244d46. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Grep
    • Glob

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

RAG Skills loads about 1.3k tokens when it runs. Until then it costs about 36 tokens; SKILL.md has 168 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~36
When it runs · the whole SKILL.md, loaded when a task matches
~1.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from llama-farm/llamafarm at commit 6244d46, republished under its Apache-2.0 licence (© llama-farm). 168 words, ~1,284 tokens.

Download SKILL.mdSave it as .claude/skills/rag-skills/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
rag-skills
description
RAG-specific best practices for LlamaIndex, ChromaDB, and Celery workers. Covers ingestion, retrieval, embeddings, and performance.
allowed-tools
Read, Grep, Glob
user-invocable
false

RAG Skills for LlamaFarm

Framework-specific patterns and code review checklists for the RAG component.

Extends: python-skills - All Python best practices apply here.

Component Overview

AspectTechnologyVersion
PythonPython3.11+
Document ProcessingLlamaIndex0.13+
Vector StorageChromaDB1.0+
Task QueueCelery5.5+
EmbeddingsUniversal/Ollama/OpenAIMultiple

Directory Structure

rag/
├── api.py                 # Search and database APIs
├── celery_app.py          # Celery configuration
├── main.py                # Entry point
├── core/
│   ├── base.py            # Document, Component, Pipeline ABCs
│   ├── factories.py       # Component factories
│   ├── ingest_handler.py  # File ingestion with safety checks
│   ├── blob_processor.py  # Binary file processing
│   ├── settings.py        # Pydantic settings
│   └── logging.py         # RAGStructLogger
├── components/
│   ├── embedders/         # Embedding providers
│   ├── extractors/        # Metadata extractors
│   ├── parsers/           # Document parsers (LlamaIndex)
│   ├── retrievers/        # Retrieval strategies
│   └── stores/            # Vector stores (ChromaDB, FAISS)
├── tasks/                 # Celery tasks
│   ├── ingest_tasks.py    # File ingestion
│   ├── search_tasks.py    # Database search
│   ├── query_tasks.py     # Complex queries
│   ├── health_tasks.py    # Health checks
│   └── stats_tasks.py     # Statistics
└── utils/
    └── embedding_safety.py  # Circuit breaker, validation

Quick Reference

TopicFileKey Points
LlamaIndexllamaindex.mdDocument parsing, chunking, node conversion
ChromaDBchromadb.mdCollections, embeddings, distance metrics
Celerycelery.mdTask routing, error handling, worker config
Performanceperformance.mdBatching, caching, deduplication

Core Patterns

Document Dataclass
python
from dataclasses import dataclass, field
from typing import Any

@dataclass
class Document:
    content: str
    metadata: dict[str, Any] = field(default_factory=dict)
    id: str = field(default_factory=lambda: str(uuid.uuid4()))
    source: str | None = None
    embeddings: list[float] | None = None
Component Abstract Base Class
python
from abc import ABC, abstractmethod

class Component(ABC):
    def __init__(
        self,
        name: str | None = None,
        config: dict[str, Any] | None = None,
        project_dir: Path | None = None,
    ):
        self.name = name or self.__class__.__name__
        self.config = config or {}
        self.logger = RAGStructLogger(__name__).bind(name=self.name)
        self.project_dir = project_dir

    @abstractmethod
    def process(self, documents: list[Document]) -> ProcessingResult:
        pass
Retrieval Strategy Pattern
python
class RetrievalStrategy(Component, ABC):
    @abstractmethod
    def retrieve(
        self,
        query_embedding: list[float],
        vector_store,
        top_k: int = 5,
        **kwargs
    ) -> RetrievalResult:
        pass

    @abstractmethod
    def supports_vector_store(self, vector_store_type: str) -> bool:
        pass
Embedder with Circuit Breaker
python
class Embedder(Component):
    DEFAULT_FAILURE_THRESHOLD = 5
    DEFAULT_RESET_TIMEOUT = 60.0

    def __init__(self, ...):
        super().__init__(...)
        self._circuit_breaker = CircuitBreaker(
            failure_threshold=config.get("failure_threshold", 5),
            reset_timeout=config.get("reset_timeout", 60.0),
        )
        self._fail_fast = config.get("fail_fast", True)

    def embed_text(self, text: str) -> list[float]:
        self.check_circuit_breaker()
        try:
            embedding = self._call_embedding_api(text)
            self.record_success()
            return embedding
        except Exception as e:
            self.record_failure(e)
            if self._fail_fast:
                raise EmbedderUnavailableError(str(e)) from e
            return [0.0] * self.get_embedding_dimension()

Review Checklist Summary

When reviewing RAG code:

  1. LlamaIndex (Medium priority)

    • Proper chunking configuration
    • Metadata preservation during parsing
    • Error handling for unsupported formats
  2. ChromaDB (High priority)

    • Thread-safe client access
    • Proper distance metric selection
    • Metadata type compatibility
  3. Celery (High priority)

    • Task routing to correct queue
    • Error logging with context
    • Proper serialization
  4. Performance (Medium priority)

    • Batch processing for embeddings
    • Deduplication enabled
    • Appropriate caching

See individual topic files for detailed checklists with grep patterns.

© llama-farm, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files in .claude/skills/rag-skills of llama-farm/llamafarm.

  • SKILL.md
  • celery.md
  • chromadb.md
  • llamaindex.md
  • performance.md

Open the folder on GitHubat commit 6244d46

Compare with similar skills

RAG Skills next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

RAG Skills compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
RAG Skills this skillllama-farm/llamafarm836—~1.3kAutomated safety check: PassApache-2.0
Chroma Vector DatabaseOrchestra-Research/AI-Research-SKILLs13k7 repos~2.3kAutomated safety check: PassMIT
Sentence Transformers EmbeddingsOrchestra-Research/AI-Research-SKILLs13k2 repos~1.6kAutomated safety check: PassMIT
FAISS Similarity SearchOrchestra-Research/AI-Research-SKILLs13k6 repos~1.3kAutomated safety check: PassMIT
Retail Product Search Agentgoogle/adk-recipes10k—~3kAutomated safety check: PassApache-2.0
Llama Index Wikichujianyun/skills742—~480Automated safety check: PassCustom licence

Similar skills

  • Chroma Vector Database

    Orchestra-Research/AI-Research-SKILLs

    Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.

    13k GitHub starsUsed in 7 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Sentence Transformers Embeddings

    Orchestra-Research/AI-Research-SKILLs

    Generates text embeddings locally with the sentence-transformers library for RAG, semantic search, clustering and similarity, with model picks for general, multilingual and legal text.

    13k GitHub starsUsed in 2 repos~1.6k tokens
    AI & LLM EngineeringAuto-check passed
  • FAISS Similarity Search

    Orchestra-Research/AI-Research-SKILLs

    Sets up FAISS for fast nearest-neighbor search over large collections of dense vectors, choosing between Flat, IVF, HNSW and product quantization indexes.

    13k GitHub starsUsed in 6 repos~1.3k tokens
    DatabasesAuto-check passed
  • Retail Product Search Agent

    google/adk-recipes

    Official

    Builds a retail product search agent on Google Cloud, from catalog ingestion into BigQuery and Vector Search to ADK scaffolding, evaluation and Cloud Run deployment.

    10k GitHub stars~3k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Llama Index Wiki

    chujianyun/skills

    LlamaIndex 官方用户文档离线知识库,用于检索并回答 LlamaIndex Python 框架的安装、RAG、数据加载、索引、检索与查询、Agent、Workflow、模型、Embedding、向量库、评估、可观测性、部署、LlamaCloud 和 LlamaParse 等问题,也可生成有文档依据的示例代码与排障建议。当用户提到…

    742 GitHub stars~480 tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Langchain RAG

    langchain-ai/langchain-skills

    Official

    INVOKE THIS SKILL when building ANY retrieval-augmented generation (RAG) system.

    1.3k GitHub stars~3.9k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed

More from llama-farm/llamafarm

All 19 skills in this repo
  • Reflect

    llama-farm/llamafarm

    Analyze the current session and propose improvements to skills.

    836 GitHub stars~1.5k tokensUpdated 4 mo ago
    Auto-check: notes
  • Temp Files

    llama-farm/llamafarm

    Guidelines for creating temporary files in system temp directory.

    836 GitHub stars~515 tokensUpdated 4 mo ago
    Auto-check: notes
  • CLI Skills

    llama-farm/llamafarm

    CLI best practices for LlamaFarm. An agent skill from llama-farm/llamafarm.

    836 GitHub stars~1.2k tokensUpdated 4 mo ago
    Auto-check passed
  • Code Review

    llama-farm/llamafarm

    Comprehensive code review for diffs. An agent skill from llama-farm/llamafarm.

    836 GitHub stars~2.3k tokensUpdated 4 mo ago
    Auto-check: notes
  • Commit Push PR

    llama-farm/llamafarm

    Commit changes, push to GitHub, and open a PR. An agent skill from llama-farm/llamafarm.

    836 GitHub stars~2.3k tokensUpdated 4 mo ago
    Auto-check: notes
  • Common Skills

    llama-farm/llamafarm

    Best practices for the Common utilities package in LlamaFarm.

    836 GitHub stars~885 tokensUpdated 4 mo ago
    Auto-check passed

Questions about RAG Skills

What does RAG Skills do?

RAG-specific best practices for LlamaIndex, ChromaDB, and Celery workers. RAG Skills is an agent skill from llama-farm/llamafarm. RAG-specific best practices for LlamaIndex, ChromaDB, and Celery workers.

When should I use RAG Skills?

RAG Skills fits situations like: tasks that involve Retrieval-augmented generation; tasks that involve Background jobs; tasks that involve Embeddings.

How do I install RAG Skills in Claude Code?

Run `npx skills add llama-farm/llamafarm --skill rag-skills -a claude-code`. Or copy the skill folder (.claude/skills/rag-skills in llama-farm/llamafarm) into .claude/skills/rag-skills in your project. Claude Code loads it when a task matches its description.

How do I install RAG Skills in Codex?

Run `npx skills add llama-farm/llamafarm --skill rag-skills -a codex`. Or copy the skill folder (.claude/skills/rag-skills in llama-farm/llamafarm) into .agents/skills/rag-skills in your project. Codex loads it when a task matches its description.

Can I use RAG Skills in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add llama-farm/llamafarm --skill rag-skills -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/rag-skills, .gemini/skills/rag-skills, .github/skills/rag-skills and .opencode/skills/rag-skills in your project.

What does RAG Skills need to run?

SKILL.md names no scripts, command-line tools or credentials: RAG Skills is instructions for the agent only. Our summary lists: Python 3. Its frontmatter pre-approves these tools: Read, Grep, Glob.

Does RAG Skills access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is RAG Skills safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does RAG Skills use?

RAG Skills is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does RAG Skills use?

About 1.3k tokens (SKILL.md is roughly 5.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to RAG Skills?

Skills that share tags, products or a category with RAG Skills: Chroma Vector Database (Orchestra-Research/AI-Research-SKILLs, 13k stars), Sentence Transformers Embeddings (Orchestra-Research/AI-Research-SKILLs, 13k stars), FAISS Similarity Search (Orchestra-Research/AI-Research-SKILLs, 13k stars) and Retail Product Search Agent (google/adk-recipes, 10k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains RAG Skills?

llama-farm (a GitHub organization) maintains it in llama-farm/llamafarm, which has 836 GitHub stars. The repository holds 19 skills in this directory. The repository was last updated on June 10, 2026.

Source: llama-farm/llamafarm on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.