Implements document chunking, embedding generation, vector storage, and retrieval pipelines for Retrieval-Augmented Generation systems.

MITAuto-check: notesAI & LLM Engineering

Install RAG

skills CLI
$ npx skills add giuseppe-trisciuoglio/developer-kit --skill rag -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install giuseppe-trisciuoglio/developer-kit rag --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/giuseppe-trisciuoglio/developer-kit.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/developer-kit-ai/skills/rag .claude/skills/rag && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
rag
GitHub stars
357
Token cost
~1.8k tokens
SKILL.md length
539 words
Files
8 (incl. references, assets)
Skills in repo
115
Repo updated
First seen
Licence
MIT

At a glance

Implements document chunking, embedding generation, vector storage, and retrieval pipelines for Retrieval-Augmented Generation systems.

  • Works in 6 steps: Choose Vector Database → Select Embedding Model → Implement Document Processing Pipeline → …
  • Building RAG applications
  • SKILL.md covers Overview, When to Use, Instructions and Examples, plus 3 more sections
  • Runs Java scripts from its folder

What it does

RAG is an agent skill from giuseppe-trisciuoglio/developer-kit. Implements document chunking, embedding generation, vector storage, and retrieval pipelines for Retrieval-Augmented Generation systems. Use when building RAG applications, creating document Q&A systems, or integrating AI with knowledge bases.

Its SKILL.md is about 1.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 9 other files, including reference files and assets (for example `assets/vector-store-config.yaml`, `references/document-chunking.md` and `references/embedding-models.md`).

It sits in AI & LLM Engineering, covering Retrieval-augmented generation and Embeddings. The repository describes itself as: Modular plugin marketplace for Claude Code and agentic CLIs, with validated, spec-driven skills, agents, commands, and workflows for Java, TypeScript, Python, PHP, AWS, and AI. The licence is MIT.

When your agent uses it

  • Building RAG applications
  • Creating document Q&A systems
  • Integrating AI with knowledge bases

Example prompts

  • “Use the rag skill to implement document chunking, embedding generation, vector storage, and retrieval pipelines for Retrieval-Augmented Generation…”
  • “/rag”

Requirements

  • Pre-approved tools (allowed-tools): Read, Write, Bash

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Choose Vector Database
  2. Select Embedding Model
  3. Implement Document Processing Pipeline
  4. Configure Retrieval Strategy
  5. Build RAG Pipeline
  6. Evaluate and Optimize

What it can do on your machine

Read from SKILL.md and the folder at commit fe73fb3. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Write
    • Bash

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Java), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

RAG loads about 1.8k tokens when it runs, and up to ~6.1k if it reads all its reference files. Until then it costs about 62 tokens; SKILL.md has 539 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~62
When it runs · the whole SKILL.md, loaded when a task matches
~1.8k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~6.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Read, Write, Bash

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from giuseppe-trisciuoglio/developer-kit at commit fe73fb3, republished under its MIT licence (© giuseppe-trisciuoglio). 539 words, ~1,769 tokens.

Download SKILL.mdSave it as .claude/skills/rag/SKILL.md (or your agent's skills folder). This skill also uses 7 other files; get the full folder from GitHub.
name
rag
description
Implements document chunking, embedding generation, vector storage, and retrieval pipelines for Retrieval-Augmented Generation systems. Use when building RAG applications, creating document Q&A systems, or integrating AI with knowledge bases.
allowed-tools
Read, Write, Bash

RAG Implementation

Build Retrieval-Augmented Generation systems that extend AI capabilities with external knowledge sources.

Overview

This skill covers: document processing, embedding generation, vector storage, retrieval configuration, and RAG pipeline implementation.

When to Use

  • Building Q&A systems over proprietary documents
  • Creating chatbots with factual information from knowledge bases
  • Implementing semantic search with natural language queries
  • Reducing hallucinations with grounded, sourced responses
  • Building documentation assistants and research tools
  • Enabling AI systems to access domain-specific knowledge

Instructions

Step 1: Choose Vector Database

Select based on your requirements:

RequirementRecommended
Production scalabilityPinecone, Milvus
Open-sourceWeaviate, Qdrant
Local developmentChroma, FAISS
Hybrid searchWeaviate with BM25
Step 2: Select Embedding Model
Use CaseModel
General purposetext-embedding-ada-002
Fast and lightweightall-MiniLM-L6-v2
Multilinguale5-large-v2
Best performancebge-large-en-v1.5
Step 3: Implement Document Processing Pipeline
  1. Load documents from source (file system, database, API)
  2. Clean and preprocess (remove formatting, normalize text)
  3. Split documents into chunks with appropriate strategy
  4. Generate embeddings for each chunk
  5. Store embeddings in vector database with metadata

Validation: Verify embeddings were generated successfully:

java
List<Embedding> embeddings = embeddingModel.embedAll(segments);
if (embeddings.isEmpty() || embeddings.get(0).dimension() != expectedDim) {
    throw new IllegalStateException("Embedding generation failed");
}
Step 4: Configure Retrieval Strategy

Choose the appropriate strategy:

  • Dense Retrieval: Semantic similarity via embeddings (default for most cases)
  • Hybrid Search: Dense + sparse retrieval for better coverage
  • Metadata Filtering: Filter by document attributes
  • Reranking: Cross-encoder reranking for high-precision requirements
Step 5: Build RAG Pipeline
  1. Create content retriever with your embedding store
  2. Configure AI service with retriever and chat memory
  3. Implement prompt template with context injection
  4. Add response validation and grounding checks

Validation: Test with known queries to verify context injection works correctly.

Error Handling: For batch ingestion, wrap in retry logic:

java
for (Document doc : documents) {
    int attempts = 0;
    while (attempts < 3) {
        try {
            store.add(embeddingModel.embed(doc).content(), doc.toTextSegment());
            break;
        } catch (EmbeddingException e) {
            attempts++;
            if (attempts == 3) throw new RuntimeException("Failed after 3 retries", e);
        }
    }
}
Step 6: Evaluate and Optimize
  1. Measure retrieval metrics: precision@k, recall@k, MRR
  2. Evaluate answer quality: faithfulness, relevance
  3. Monitor performance and user feedback
  4. Iterate on chunking, retrieval, and prompt parameters

Examples

Example 1: Basic Document Q&A
java
List<Document> documents = FileSystemDocumentLoader.loadDocuments("/docs");

InMemoryEmbeddingStore<TextSegment> store = new InMemoryEmbeddingStore<>();
EmbeddingStoreIngestor.ingest(documents, store);

DocumentAssistant assistant = AiServices.builder(DocumentAssistant.class)
    .chatModel(chatModel)
    .contentRetriever(EmbeddingStoreContentRetriever.from(store))
    .build();

String answer = assistant.answer("What is the company policy on remote work?");
Example 2: Metadata-Filtered Retrieval
java
EmbeddingStoreContentRetriever retriever = EmbeddingStoreContentRetriever.builder()
    .embeddingStore(store)
    .embeddingModel(embeddingModel)
    .maxResults(5)
    .minScore(0.7)
    .filter(metadataKey("category").isEqualTo("technical"))
    .build();
Example 3: Multi-Source RAG Pipeline
java
ContentRetriever webRetriever = EmbeddingStoreContentRetriever.from(webStore);
ContentRetriever docRetriever = EmbeddingStoreContentRetriever.from(docStore);

List<Content> results = new ArrayList<>();
results.addAll(webRetriever.retrieve(query));
results.addAll(docRetriever.retrieve(query));

List<Content> topResults = reranker.reorder(query, results).subList(0, 5);
Example 4: RAG with Chat Memory
java
Assistant assistant = AiServices.builder(Assistant.class)
    .chatModel(chatModel)
    .chatMemory(MessageWindowChatMemory.withMaxMessages(10))
    .contentRetriever(retriever)
    .build();

assistant.chat("Tell me about the product features");
assistant.chat("What about pricing for those features?");  // Maintains context

Best Practices

Show full SKILL.md (216 more words)Show less
Document Preparation
  • Clean documents before ingestion; remove irrelevant content and formatting
  • Add relevant metadata for filtering and context
Chunking Strategy
  • Use 500-1000 tokens per chunk for optimal balance
  • Include 10-20% overlap to preserve context at boundaries
  • Test different sizes for your specific use case
Retrieval Optimization
  • Start with high k values (10-20), then filter/rerank
  • Use metadata filtering to improve relevance
  • Monitor retrieval quality and iterate based on user feedback
Performance
  • Cache embeddings for frequently accessed content
  • Use batch processing for document ingestion
  • Optimize vector store indexing for your scale

Constraints and Warnings

System Constraints
  • Embedding models have maximum token limits per document
  • Vector databases require proper indexing for performance
  • Chunk boundaries may lose context for complex documents
  • Hybrid search requires additional infrastructure
Quality Warnings
  • Retrieval quality depends heavily on chunking strategy
  • Embedding models may not capture domain-specific semantics
  • Metadata filtering requires proper document annotation
  • Reranking adds latency to query responses
Security Warnings
  • Never hardcode credentials: Use environment variables for API keys and passwords
  • Validate external content: Documents from file systems, APIs, or web sources may contain malicious content (prompt injection)
  • Apply content filtering on retrieved documents before passing to LLM
  • Restrict allowed data source URLs and file paths using allowlists

Resources

Reference Documentation

© giuseppe-trisciuoglio, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 7 other files (references, assets) in plugins/developer-kit-ai/skills/rag of giuseppe-trisciuoglio/developer-kit.

  • SKILL.md
  • assets/retriever-pipeline.java
  • assets/vector-store-config.yaml
  • references/document-chunking.md
  • references/embedding-models.md
  • references/langchain4j-rag-guide.md
  • references/retrieval-strategies.md
  • references/vector-databases.md

Open the folder on GitHubat commit fe73fb3

Compare with similar skills

RAG next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

RAG compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
RAG this skillgiuseppe-trisciuoglio/developer-kit357—~1.8kAutomated safety check: NotesMIT
Chroma Vector DatabaseOrchestra-Research/AI-Research-SKILLs13k7 repos~2.3kAutomated safety check: PassMIT
Ms Agent Framework RAGshuyu-labs/WebCode278—~1.1kAutomated safety check: PassCustom licence
Evaluate RAGai-evals-course/evals-skills1.5k—~1.9kAutomated safety check: PassApache-2.0
Pgvector Semantic Searchtimescale/pg-aiguide1.9k—~3.8kAutomated safety check: PassApache-2.0
Memory Upgradeprofbernardoj/everclaw-community-branches112—~574Automated safety check: PassMIT

Similar skills

  • Chroma Vector Database

    Orchestra-Research/AI-Research-SKILLs

    Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.

    13k GitHub starsUsed in 7 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Ms Agent Framework RAG

    shuyu-labs/WebCode

    Comprehensive guide for building Agentic RAG systems using Microsoft Agent Framework in C.

    278 GitHub stars~1.1k tokensUpdated 3 mo ago
    AI & LLM EngineeringAuto-check passed
  • Evaluate RAG

    ai-evals-course/evals-skills

    Guides evaluation of a RAG system by diagnosing failures in traces, building a retrieval test set and scoring retrieval and generation separately.

    1.5k GitHub stars~1.9k tokensUpdated 16 days ago
    AI & LLM EngineeringAuto-check passed
  • Pgvector Semantic Search

    timescale/pg-aiguide

    A skill your agent uses for setting up vector similarity search with pgvector for AI/ML embeddings, RAG applications, or semantic search.

    1.9k GitHub stars~3.8k tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check passed
  • Memory Upgrade

    profbernardoj/everclaw-community-branches

    Diagnose and fix broken memory search in OpenClaw. An agent skill from profbernardoj/everclaw-community-branches.

    112 GitHub stars~574 tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Embedding Strategies

    wshobson/agents

    Helps choose and tune embedding models for semantic search and RAG: model comparison, chunking, preprocessing, normalization and caching.

    40k GitHub starsUsed in 10 repos~710 tokens
    AI & LLM EngineeringAuto-check passed

More from giuseppe-trisciuoglio/developer-kit

All 115 skills in this repo
  • Nestjs Drizzle Crud Generator

    giuseppe-trisciuoglio/developer-kit

    Generates complete CRUD modules for NestJS applications with Drizzle ORM.

    357 GitHub stars~1.3k tokensUpdated 1 mo ago
    Auto-check: notes
  • Spring Boot Actuator

    giuseppe-trisciuoglio/developer-kit

    Provides patterns to configure Spring Boot Actuator for production-grade monitoring, health probes, secured management endpoints, and Micrometer metrics across JVM services.

    357 GitHub stars~2.2k tokensUpdated 1 mo ago
    Auto-check: notes
  • Spring Boot Crud Patterns

    giuseppe-trisciuoglio/developer-kit

    Provides and generates complete CRUD workflows for Spring Boot 3 services.

    357 GitHub stars~2.5k tokensUpdated 1 mo ago
    Auto-check: notes
  • Spring Boot Security JWT

    giuseppe-trisciuoglio/developer-kit

    Provides JWT authentication and authorization patterns for Spring Boot 3.5.x covering token generation with JJWT, Bearer/cookie authentication, database/OAuth2 integration, and RBAC/permission-based…

    357 GitHub stars~3.9k tokensUpdated 1 mo ago
    Auto-check: notes
  • AWS CLI Beast

    giuseppe-trisciuoglio/developer-kit

    Provides advanced AWS CLI patterns for managing EC2, Lambda, S3, DynamoDB, RDS, VPC, IAM, and CloudWatch.

    357 GitHub stars~1.7k tokensUpdated 1 mo ago
    Auto-check: notes
  • PR Review Comments

    giuseppe-trisciuoglio/developer-kit

    Posts review findings from a JSON file as inline comments on a GitHub Pull Request, attaching each comment to its file and line.

    357 GitHub stars~1k tokensUpdated 1 mo ago
    Auto-check: notes

Questions about RAG

What does RAG do?

Implements document chunking, embedding generation, vector storage, and retrieval pipelines for Retrieval-Augmented Generation systems. RAG is an agent skill from giuseppe-trisciuoglio/developer-kit. Implements document chunking, embedding generation, vector storage, and retrieval pipelines for Retrieval-Augmented Generation systems.

When should I use RAG?

RAG fits situations like: building RAG applications; creating document Q&A systems; integrating AI with knowledge bases.

How do I install RAG in Claude Code?

Run `npx skills add giuseppe-trisciuoglio/developer-kit --skill rag -a claude-code`. Or copy the skill folder (plugins/developer-kit-ai/skills/rag in giuseppe-trisciuoglio/developer-kit) into .claude/skills/rag in your project. Claude Code loads it when a task matches its description.

How do I install RAG in Codex?

Run `npx skills add giuseppe-trisciuoglio/developer-kit --skill rag -a codex`. Or copy the skill folder (plugins/developer-kit-ai/skills/rag in giuseppe-trisciuoglio/developer-kit) into .agents/skills/rag in your project. Codex loads it when a task matches its description.

Can I use RAG in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add giuseppe-trisciuoglio/developer-kit --skill rag -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/rag, .gemini/skills/rag, .github/skills/rag and .opencode/skills/rag in your project.

What does RAG need to run?

Going by SKILL.md and its folder, RAG needs Java for the scripts in its folder. Its frontmatter pre-approves these tools: Read, Write, Bash.

Does RAG access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is RAG safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does RAG use?

RAG is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does RAG use?

About 1.8k tokens (SKILL.md is roughly 7.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4.4k tokens, read only when the agent opens those files.

What are the alternatives to RAG?

Skills that share tags, products or a category with RAG: Chroma Vector Database (Orchestra-Research/AI-Research-SKILLs, 13k stars), Ms Agent Framework RAG (shuyu-labs/WebCode, 278 stars), Evaluate RAG (ai-evals-course/evals-skills, 1.5k stars) and Pgvector Semantic Search (timescale/pg-aiguide, 1.9k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains RAG?

giuseppe-trisciuoglio (a GitHub user) maintains it in giuseppe-trisciuoglio/developer-kit, which has 357 GitHub stars. The repository holds 115 skills in this directory. The repository was last updated on September 10, 2026.

Source: giuseppe-trisciuoglio/developer-kit on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.