Agent skill

RAG Engineer

by majiayu000 in majiayu000/claude-skill-registry

Expert in building Retrieval-Augmented Generation systems. An agent skill from majiayu000/claude-skill-registry.

MITAuto-check passedAI & LLM Engineering

Install RAG Engineer

skills CLI
$ npx skills add majiayu000/claude-skill-registry --skill rag-engineer -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install majiayu000/claude-skill-registry rag-engineer --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/majiayu000/claude-skill-registry.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/ai-llm/rag-engineer-sickn33-antigravity-awesome .claude/skills/rag-engineer && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
rag-engineer
GitHub stars
666
Used in
3 other repos
Token cost
~2.5k tokens
SKILL.md length
1,237 words
Files
2
Skills in repo
971
Repo updated
First seen
Licence
MIT

At a glance

Expert in building Retrieval-Augmented Generation systems. An agent skill from majiayu000/claude-skill-registry.

  • Tasks that involve Retrieval-augmented generation
  • SKILL.md covers Capabilities, Prerequisites, Patterns and Sharp Edges, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Tasks that involve Embeddings

What it does

RAG Engineer is an agent skill from majiayu000/claude-skill-registry. Expert in building Retrieval-Augmented Generation systems. Masters embedding models, vector databases, chunking strategies, and retrieval optimization for LLM applications.

Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `metadata.json`).

It sits in AI & LLM Engineering, covering Retrieval-augmented generation, Embeddings and Vector databases. The repository describes itself as: Searchable Claude Code skills catalog with source-linked guides and generated registry artifacts. The licence is MIT.

When your agent uses it

  • Tasks that involve Retrieval-augmented generation
  • Tasks that involve Embeddings
  • Tasks that involve Vector databases

Example prompts

  • “/rag-engineer”

What it can do on your machine

Read from SKILL.md and the folder at commit 000116a. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

RAG Engineer loads about 2.5k tokens when it runs. Until then it costs about 46 tokens; SKILL.md has 1,237 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~46
When it runs · the whole SKILL.md, loaded when a task matches
~2.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from majiayu000/claude-skill-registry at commit 000116a, republished under its MIT licence (© majiayu000). 1,237 words, ~2,489 tokens.

Download SKILL.mdSave it as .claude/skills/rag-engineer/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
rag-engineer
description
Expert in building Retrieval-Augmented Generation systems. Masters embedding models, vector databases, chunking strategies, and retrieval optimization for LLM applications.
risk
unknown
source
vibeship-spawner-skills (Apache 2.0)
date_added
2026-02-27

RAG Engineer

Expert in building Retrieval-Augmented Generation systems. Masters embedding models, vector databases, chunking strategies, and retrieval optimization for LLM applications.

Role: RAG Systems Architect

I bridge the gap between raw documents and LLM understanding. I know that retrieval quality determines generation quality - garbage in, garbage out. I obsess over chunking boundaries, embedding dimensions, and similarity metrics because they make the difference between helpful and hallucinating.

Expertise
  • Embedding model selection and fine-tuning
  • Vector database architecture and scaling
  • Chunking strategies for different content types
  • Retrieval quality optimization
  • Hybrid search implementation
  • Re-ranking and filtering strategies
  • Context window management
  • Evaluation metrics for retrieval
Principles
  • Retrieval quality > Generation quality - fix retrieval first
  • Chunk size depends on content type and query patterns
  • Embeddings are not magic - they have blind spots
  • Always evaluate retrieval separately from generation
  • Hybrid search beats pure semantic in most cases

Capabilities

  • Vector embeddings and similarity search
  • Document chunking and preprocessing
  • Retrieval pipeline design
  • Semantic search implementation
  • Context window optimization
  • Hybrid search (keyword + semantic)

Prerequisites

  • Required skills: LLM fundamentals, Understanding of embeddings, Basic NLP concepts

Patterns

Semantic Chunking

Chunk by meaning, not arbitrary token counts

When to use: Processing documents with natural sections

  • Use sentence boundaries, not token limits
  • Detect topic shifts with embedding similarity
  • Preserve document structure (headers, paragraphs)
  • Include overlap for context continuity
  • Add metadata for filtering
Hierarchical Retrieval

Multi-level retrieval for better precision

When to use: Large document collections with varied granularity

  • Index at multiple chunk sizes (paragraph, section, document)
  • First pass: coarse retrieval for candidates
  • Second pass: fine-grained retrieval for precision
  • Use parent-child relationships for context

Combine semantic and keyword search

When to use: Queries may be keyword-heavy or semantic

  • BM25/TF-IDF for keyword matching
  • Vector similarity for semantic matching
  • Reciprocal Rank Fusion for combining scores
  • Weight tuning based on query type
Query Expansion

Expand queries to improve recall

When to use: User queries are short or ambiguous

  • Use LLM to generate query variations
  • Add synonyms and related terms
  • Hypothetical Document Embedding (HyDE)
  • Multi-query retrieval with deduplication
Contextual Compression

Compress retrieved context to fit window

When to use: Retrieved chunks exceed context limits

  • Extract relevant sentences only
  • Use LLM to summarize chunks
  • Remove redundant information
  • Prioritize by relevance score
Metadata Filtering

Pre-filter by metadata before semantic search

When to use: Documents have structured metadata

  • Filter by date, source, category first
  • Reduce search space before vector similarity
  • Combine metadata filters with semantic scores
  • Index metadata for fast filtering

Sharp Edges

Fixed-size chunking breaks sentences and context

Severity: HIGH

Situation: Using fixed token/character limits for chunking

Symptoms:

  • Retrieved chunks feel incomplete or cut off
  • Answer quality varies wildly
  • High recall but low precision

Why this breaks: Fixed-size chunks split mid-sentence, mid-paragraph, or mid-idea. The resulting embeddings represent incomplete thoughts, leading to poor retrieval quality. Users search for concepts but get fragments.

Recommended fix:

Use semantic chunking that respects document structure:

  • Split on sentence/paragraph boundaries
  • Use embedding similarity to detect topic shifts
  • Include overlap for context continuity
  • Preserve headers and document structure as metadata
Pure semantic search without metadata pre-filtering

Severity: MEDIUM

Situation: Only using vector similarity, ignoring metadata

Symptoms:

  • Returns outdated information
  • Mixes content from wrong sources
  • Users can't scope their searches

Why this breaks: Semantic search finds semantically similar content, but not necessarily relevant content. Without metadata filtering, you return old docs when user wants recent, wrong categories, or inapplicable content.

Recommended fix:

Implement hybrid filtering:

  • Pre-filter by metadata (date, source, category) before vector search
  • Post-filter results by relevance criteria
  • Include metadata in the retrieval API
  • Allow users to specify filters
Using same embedding model for different content types

Severity: MEDIUM

Situation: One embedding model for code, docs, and structured data

Symptoms:

  • Code search returns irrelevant results
  • Domain terms not matched properly
  • Similar concepts not clustered

Why this breaks: Embedding models are trained on specific content types. Using a text embedding model for code, or a general model for domain-specific content, produces poor similarity matches.

Recommended fix:

Evaluate embeddings per content type:

  • Use code-specific embeddings for code (e.g., CodeBERT)
  • Consider domain-specific or fine-tuned embeddings
  • Benchmark retrieval quality before choosing
  • Separate indices for different content types if needed
Using first-stage retrieval results directly

Severity: MEDIUM

Situation: Taking top-K from vector search without reranking

Symptoms:

  • Clearly relevant docs not in top results
  • Results order seems arbitrary
  • Adding more results helps quality

Why this breaks: First-stage retrieval (vector search) optimizes for recall, not precision. The top results by embedding similarity may not be the most relevant for the specific query. Cross-encoder reranking dramatically improves precision for the final results.

Recommended fix:

Add reranking step:

  • Retrieve larger candidate set (e.g., top 20-50)
  • Rerank with cross-encoder (query-document pairs)
  • Return reranked top-K (e.g., top 5)
  • Cache reranker for performance
Show full SKILL.md (462 more words)Show less
Cramming maximum context into LLM prompt

Severity: MEDIUM

Situation: Using all retrieved context regardless of relevance

Symptoms:

  • Answers drift with more context
  • LLM ignores key information
  • High token costs

Why this breaks: More context isn't always better. Irrelevant context confuses the LLM, increases latency and cost, and can cause the model to ignore the most relevant information. Models have attention limits.

Recommended fix:

Use relevance thresholds:

  • Set minimum similarity score cutoff
  • Limit context to truly relevant chunks
  • Summarize or compress if needed
  • Order context by relevance
Not measuring retrieval quality separately from generation

Severity: HIGH

Situation: Only evaluating end-to-end RAG quality

Symptoms:

  • Can't diagnose poor RAG performance
  • Prompt changes don't help
  • Random quality variations

Why this breaks: If answers are wrong, you can't tell if retrieval failed or generation failed. This makes debugging impossible and leads to wrong fixes (tuning prompts when retrieval is the problem).

Recommended fix:

Separate retrieval evaluation:

  • Create retrieval test set with relevant docs labeled
  • Measure MRR, NDCG, Recall@K for retrieval
  • Evaluate generation only on correct retrievals
  • Track metrics over time
Not updating embeddings when source documents change

Severity: MEDIUM

Situation: Embeddings generated once, never refreshed

Symptoms:

  • Returns outdated information
  • References deleted content
  • Inconsistent with source

Why this breaks: Documents change but embeddings don't. Users retrieve outdated content or, worse, content that no longer exists. This erodes trust in the system.

Recommended fix:

Implement embedding refresh:

  • Track document versions/hashes
  • Re-embed on document change
  • Handle deleted documents
  • Consider TTL for embeddings
Same retrieval strategy for all query types

Severity: MEDIUM

Situation: Using pure semantic search for keyword-heavy queries

Symptoms:

  • Exact term searches miss results
  • Concept searches too literal
  • Users frustrated with both

Why this breaks: Some queries are keyword-oriented (looking for specific terms) while others are semantic (looking for concepts). Pure semantic search fails on exact matches; pure keyword search fails on paraphrases.

Recommended fix:

Implement hybrid search:

  • BM25/TF-IDF for keyword matching
  • Vector similarity for semantic matching
  • Reciprocal Rank Fusion to combine
  • Tune weights based on query patterns

Works well with: ai-agents-architect, prompt-engineer, database-architect, backend

When to Use

  • User mentions or implies: building RAG
  • User mentions or implies: vector search
  • User mentions or implies: embeddings
  • User mentions or implies: semantic search
  • User mentions or implies: document retrieval
  • User mentions or implies: context retrieval
  • User mentions or implies: knowledge base
  • User mentions or implies: LLM with documents
  • User mentions or implies: chunking strategy
  • User mentions or implies: pinecone
  • User mentions or implies: weaviate
  • User mentions or implies: chromadb
  • User mentions or implies: pgvector

Limitations

  • Use this skill only when the task clearly matches the scope described above.
  • Do not treat the output as a substitute for environment-specific validation, testing, or expert review.
  • Stop and ask for clarification if required inputs, permissions, safety boundaries, or success criteria are missing.

© majiayu000, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in skills/ai-llm/rag-engineer-sickn33-antigravity-awesome of majiayu000/claude-skill-registry.

  • SKILL.md
  • metadata.json

Open the folder on GitHubat commit 000116a

Used in 3 other repositories

We found 16 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 3 other GitHub owners. This page covers the copy in majiayu000/claude-skill-registry, which our catalogue first saw on October 7, 2026.

Compare with similar skills

RAG Engineer next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

RAG Engineer compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
RAG Engineer this skillmajiayu000/claude-skill-registry6663 repos~2.5kAutomated safety check: PassMIT
Chroma Vector DatabaseOrchestra-Research/AI-Research-SKILLs13k8 repos~2.3kAutomated safety check: PassMIT
Ms Agent Framework RAGshuyu-labs/WebCode278—~1.1kAutomated safety check: PassCustom licence
Pgvector Semantic Searchtimescale/pg-aiguide1.9k1 repos~3.8kAutomated safety check: PassApache-2.0
RAG ArchitectJeffallan/claude-skills12k1 repos~2kAutomated safety check: PassMIT
RAG Implementationwshobson/agents40k9 repos~1.1kAutomated safety check: PassMIT

Similar skills

  • Chroma Vector Database

    Orchestra-Research/AI-Research-SKILLs

    Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.

    13k GitHub starsUsed in 8 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Ms Agent Framework RAG

    shuyu-labs/WebCode

    Comprehensive guide for building Agentic RAG systems using Microsoft Agent Framework in C.

    278 GitHub stars~1.1k tokensUpdated 3 mo ago
    AI & LLM EngineeringAuto-check passed
  • Pgvector Semantic Search

    timescale/pg-aiguide

    A skill your agent uses for setting up vector similarity search with pgvector for AI/ML embeddings, RAG applications, or semantic search.

    1.9k GitHub starsUsed in 1 repo~3.8k tokens
    AI & LLM EngineeringAuto-check passed
  • RAG Architect

    Jeffallan/claude-skills

    Designs retrieval-augmented generation systems: document chunking, embeddings, vector store setup, hybrid search, reranking and retrieval evaluation, with checks at each step.

    12k GitHub starsUsed in 1 repo~2k tokens
    AI & LLM EngineeringAuto-check passed
  • RAG Implementation

    wshobson/agents

    Build retrieval-augmented generation systems: pick a vector database and embedding model, choose retrieval and reranking strategies, and start from a LangGraph pipeline.

    40k GitHub starsUsed in 9 repos~1.1k tokens
    AI & LLM EngineeringAuto-check passed
  • RAG Engineer

    davila7/claude-code-templates

    Expert in building Retrieval-Augmented Generation systems. An agent skill from davila7/claude-code-templates.

    32k GitHub starsUsed in 5 repos~729 tokens
    AI & LLM EngineeringAuto-check passed

More from majiayu000/claude-skill-registry

All 971 skills in this repo
  • Deep Research

    majiayu000/claude-skill-registry

    Multi-source deep research using firecrawl and exa MCPs. An agent skill from majiayu000/claude-skill-registry.

    666 GitHub starsUsed in 6 repos~1.1k tokens
    Auto-check passed
  • Exa Search

    majiayu000/claude-skill-registry

    Neural search via Exa MCP for web, code, and company research.

    666 GitHub starsUsed in 5 repos~856 tokens
    Auto-check passed
  • Fal AI Media

    majiayu000/claude-skill-registry

    Unified media generation via fal.ai MCP — image, video, and audio.

    666 GitHub starsUsed in 5 repos~1.7k tokens
    Auto-check passed
  • Bgpt Paper Search

    majiayu000/claude-skill-registry

    Search scientific papers and retrieve structured experimental data extracted from full-text studies via the BGPT MCP server.

    666 GitHub starsUsed in 4 repos~619 tokens
    Auto-check: notes
  • Bio Alignment Pairwise

    majiayu000/claude-skill-registry

    Perform pairwise sequence alignment using Biopython Bio.Align.PairwiseAligner.

    666 GitHub starsUsed in 4 repos~1.7k tokens
    Auto-check passed
  • Open Notebook

    majiayu000/claude-skill-registry

    Self-hosted, open-source alternative to Google NotebookLM for AI-powered research and document analysis.

    666 GitHub starsUsed in 4 repos~2.4k tokens
    Auto-check passed

Questions about RAG Engineer

What does RAG Engineer do?

Expert in building Retrieval-Augmented Generation systems. An agent skill from majiayu000/claude-skill-registry. RAG Engineer is an agent skill from majiayu000/claude-skill-registry. Expert in building Retrieval-Augmented Generation systems.

When should I use RAG Engineer?

RAG Engineer fits situations like: tasks that involve Retrieval-augmented generation; tasks that involve Embeddings; tasks that involve Vector databases.

How do I install RAG Engineer in Claude Code?

Run `npx skills add majiayu000/claude-skill-registry --skill rag-engineer -a claude-code`. Or copy the skill folder (skills/ai-llm/rag-engineer-sickn33-antigravity-awesome in majiayu000/claude-skill-registry) into .claude/skills/rag-engineer in your project. Claude Code loads it when a task matches its description.

How do I install RAG Engineer in Codex?

Run `npx skills add majiayu000/claude-skill-registry --skill rag-engineer -a codex`. Or copy the skill folder (skills/ai-llm/rag-engineer-sickn33-antigravity-awesome in majiayu000/claude-skill-registry) into .agents/skills/rag-engineer in your project. Codex loads it when a task matches its description.

Can I use RAG Engineer in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add majiayu000/claude-skill-registry --skill rag-engineer -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/rag-engineer, .gemini/skills/rag-engineer, .github/skills/rag-engineer and .opencode/skills/rag-engineer in your project.

What does RAG Engineer need to run?

SKILL.md names no scripts, command-line tools or credentials: RAG Engineer is instructions for the agent only.

Does RAG Engineer access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is RAG Engineer safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does RAG Engineer use?

RAG Engineer is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does RAG Engineer use?

About 2.5k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to RAG Engineer?

Skills that share tags, products or a category with RAG Engineer: Chroma Vector Database (Orchestra-Research/AI-Research-SKILLs, 13k stars), Ms Agent Framework RAG (shuyu-labs/WebCode, 278 stars), Pgvector Semantic Search (timescale/pg-aiguide, 1.9k stars) and RAG Architect (Jeffallan/claude-skills, 12k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains RAG Engineer?

majiayu000 (a GitHub user) maintains it in majiayu000/claude-skill-registry, which has 666 GitHub stars. The repository holds 971 skills in this directory. The repository was last updated on October 7, 2026.

Source: majiayu000/claude-skill-registry on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.