Agent skill

RAG Engineer

by diegosouzapw in diegosouzapw/awesome-omni-skills

RAG workflow skill. An agent skill from diegosouzapw/awesome-omni-skills.

MITAuto-check passedAI & LLM Engineering

Install RAG Engineer

skills CLI
$ npx skills add diegosouzapw/awesome-omni-skills --skill rag-engineer -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install diegosouzapw/awesome-omni-skills rag-engineer --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/diegosouzapw/awesome-omni-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills_omni/rag-engineer .claude/skills/rag-engineer && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
rag-engineer
GitHub stars
159
Token cost
~3.5k tokens
SKILL.md length
1,692 words
Files
9 (incl. scripts, references)
Skills in repo
39
Repo updated
First seen
Licence
MIT

At a glance

RAG workflow skill. An agent skill from diegosouzapw/awesome-omni-skills.

  • Works in 10 steps: Define the retrieval job before… → Build a small retrieval eval set first → Define the ingestion contract → …
  • A user needs retrieval pipelines
  • SKILL.md covers Overview, When to Use This Skill, Operating Table and Workflow, plus 5 more sections
  • Runs Python scripts from its folder

What it does

RAG Engineer is an agent skill from diegosouzapw/awesome-omni-skills. RAG workflow skill. Use this skill when a user needs retrieval pipelines, chunking, ranking, citations, and evaluation for an AI application.

Its SKILL.md is about 3.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 12 other files, including scripts and reference files (for example `ATTRIBUTION.md`, `OMNI_ENHANCED.json` and `agents/openai.yaml`).

It sits in AI & LLM Engineering, covering Retrieval-augmented generation. The repository describes itself as: Public repository of AI coding skills, curated improved best-practice skills, and runtime surfaces for CLI, API, MCP, and A2A. The licence is MIT.

When your agent uses it

  • A user needs retrieval pipelines
  • Evaluation for an AI application

Example prompts

  • “/rag-engineer”

Requirements

  • Python 3

Workflow steps

10 steps, taken from the step headings in SKILL.md.

  1. Define the retrieval job before selecting tools
  2. Build a small retrieval eval set first
  3. Define the ingestion contract
  4. Choose chunking based on content structure and queries
  5. Choose retrieval strategy for the corpus
  6. Instrument the pipeline
  7. Evaluate retrieval separately from answer generation
  8. Tune in the right order
  9. Budget latency and cost by stage
  10. Define production acceptance criteria

What it can do on your machine

Read from SKILL.md and the folder at commit c3af004. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • platform.openai.com
    • developers.openai.com
    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

RAG Engineer loads about 3.5k tokens when it runs, and up to ~5.3k if it reads all its reference files. Until then it costs about 39 tokens; SKILL.md has 1,692 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~39
When it runs · the whole SKILL.md, loaded when a task matches
~3.5k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~5.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from diegosouzapw/awesome-omni-skills at commit c3af004, republished under its MIT licence (© diegosouzapw). 1,692 words, ~3,517 tokens.

Download SKILL.mdSave it as .claude/skills/rag-engineer/SKILL.md (or your agent's skills folder). This skill also uses 8 other files; get the full folder from GitHub.
name
rag-engineer
description
RAG workflow skill. Use this skill when a user needs retrieval pipelines, chunking, ranking, citations, and evaluation for an AI application.
version
0.0.1
category
ai-agents
tags
rag, retrieval, embeddings, evaluation, citations, indexing, omni-enhanced
complexity
advanced
risk
safe
tools
claude-code, cursor, gemini-cli, codex-cli, opencode
source
omni-team
author
Omni Skills Team
date_added
2026-03-27
date_updated
2026-04-19

RAG Engineer

Overview

Use this skill when the user needs a Retrieval-Augmented Generation workflow that is measurable, debuggable, and grounded in evidence.

This skill is for designing or improving:

  • corpus preparation and ingestion
  • chunking and metadata strategy
  • embedding and indexing choices
  • semantic, keyword, or hybrid retrieval
  • reranking and context assembly
  • citation and provenance behavior
  • retrieval evaluation and troubleshooting

The operating principle is simple: fix retrieval before tuning generation. If the right evidence is not found, ranked, filtered, and assembled correctly, prompt changes will mostly mask the problem.

Use the companion files when needed:

  • references/domain-notes.md for chunking decisions, hybrid retrieval rules, metrics, and failure lookup
  • examples/worked-example.md for a concrete end-to-end RAG tuning example

When to Use This Skill

Use this skill when:

  • the user is building or repairing a knowledge-grounded assistant, search-backed chat system, or internal question-answering workflow
  • the user needs help with embeddings, vector search, chunking, indexing, hybrid retrieval, reranking, or citations
  • the user needs a retrieval eval plan instead of prompt-only iteration
  • the corpus contains documents where provenance, freshness, filtering, or permissions matter
  • the system must explain which retrieved evidence supports an answer

Do not use this skill by itself when:

  • the task is mainly model fine-tuning without a retrieval component
  • the task is generic web search UX rather than document-grounded retrieval engineering
  • document permissions, tenant boundaries, or provenance cannot be enforced
  • the user only wants a one-off prompt and there is no retrieval pipeline to design or debug

Operating Table

SituationStart hereWhy it mattersMinimum acceptable outcome
New RAG systemDefine corpus slices, users, and query classesPrevents building retrieval with no target behaviorNamed corpus scope and at least 3 realistic query classes
Existing system gives weak answersCheck retrieval metrics before promptsBad retrieval often looks like bad generationA small eval set with expected supporting passages
Chunking designUse references/domain-notes.md chunking matrixChunking should follow document structure and query behaviorChunks preserve boundaries and carry useful metadata
Identifier-heavy corpusTest hybrid retrieval, not semantic-onlyError codes, version strings, SKUs, and policy numbers are easy to miss semanticallyKeyword or metadata path validated on identifier queries
Security-sensitive corpusDesign ACL and tenant filtering firstRetrieval can leak data even if generation is safeAuthorization filters applied before or during retrieval
Production tuningSet stage budgets for retrieve, rerank, assemble, answerLatency and cost failures often come from over-retrievingBudget recorded per stage with at least one trimming plan
Debugging failuresUse troubleshooting section plus references/domain-notes.mdFast diagnosis depends on mapping symptoms to pipeline stagesA suspected failure mode tied to evidence from logs or evals
Team handoffRecord corpus version, metadata schema, eval set, and known limitsMakes retrieval behavior reproducibleAnother operator can rerun the same checks

Workflow

1. Define the retrieval job before selecting tools

Document:

  • what corpus or corpora are in scope
  • who is allowed to retrieve which data
  • what query classes matter most
  • what a good retrieval result looks like
  • what downstream answer behavior is required

At minimum, identify query classes such as:

  • factual lookup
  • semantic paraphrase
  • identifier lookup
  • policy or compliance lookup
  • recent or freshness-sensitive lookup
  • multi-hop or comparison queries

Do not start with embedding model or vector database debates. Start with expected retrieval behavior.

2. Build a small retrieval eval set first

Before tuning chunk size, prompts, or ranking:

  • collect representative queries
  • record the expected supporting document or passage for each
  • separate retrieval success from answer quality
  • keep the set small but realistic so it can be rerun often

Useful eval artifacts:

  • query text
  • query class
  • expected document IDs or passage IDs
  • any required metadata filters
  • notes on ambiguity or acceptable alternatives

If the team cannot agree on expected evidence for a query, the requirement is probably underspecified.

3. Define the ingestion contract

Specify how documents become retrievable records:

  • normalization rules
  • deduplication rules
  • document identifiers
  • section extraction rules
  • freshness fields
  • ACL or tenant metadata
  • source URL, file path, title, timestamp, version, and provenance fields

Good ingestion contracts make debugging possible later. Every chunk should be traceable back to a source document and section.

4. Choose chunking based on content structure and queries

Chunk by document-aware boundaries where possible, not by arbitrary length alone.

Preserve metadata that supports retrieval and citations:

  • document ID
  • section heading or path
  • timestamp or effective date
  • source type
  • ACL or tenant tags
  • version or freshness markers

Use references/domain-notes.md for a content-type chunking matrix and common failure patterns.

Avoid assuming one universal chunk size, overlap, or top-k value. Treat these as testable starting points, not truths.

5. Choose retrieval strategy for the corpus

Select retrieval behavior that matches the corpus:

  • semantic retrieval for concept-heavy natural language content
  • keyword or lexical retrieval for identifiers, exact phrases, version strings, and error codes
  • metadata filtering for access control, tenant isolation, time ranges, product families, or content type
  • hybrid retrieval when both semantic similarity and exact matching matter
  • reranking when first-pass retrieval has adequate recall but poor ordering

A practical default is to test:

  1. semantic-only
  2. keyword-only or lexical fallback for exact terms
  3. hybrid retrieval
  4. hybrid plus reranking if latency and cost allow
6. Instrument the pipeline

Make the system observable enough to answer:

  • which chunks were retrieved
  • with what scores or rank positions
  • under which filters
  • from which corpus version
  • which chunks were passed to the model
  • which citations appeared in the answer

Log safely. Do not leak restricted content in debug traces. If needed, log chunk IDs and metadata instead of full text.

7. Evaluate retrieval separately from answer generation

Run retrieval checks before changing prompts.

For each query, ask:

  • Was the relevant document retrieved at all?
  • Was it retrieved high enough to survive truncation or reranking?
  • Did filters wrongly exclude it?
  • Did duplicate or near-duplicate chunks crowd out diversity?
  • Did context assembly omit the best evidence?

Then evaluate answer behavior separately:

  • correct use of retrieved evidence
  • citation correctness
  • abstention when evidence is weak
  • handling of ambiguity or missing context

Do not blur retrieval failure with answer synthesis failure.

Show full SKILL.md (697 more words)Show less
8. Tune in the right order

Preferred tuning order:

  1. corpus scope and data quality
  2. ingestion and deduplication
  3. metadata schema and filters
  4. chunking strategy
  5. retrieval method
  6. reranking
  7. context assembly
  8. answer prompt and response policy

This order prevents prompt work from hiding broken retrieval.

9. Budget latency and cost by stage

Track major stages such as:

  • ingest and embedding generation
  • first-pass retrieval
  • reranking
  • context assembly
  • final answer generation

If the system is slow or expensive, trim in this order first:

  • remove unnecessary retrieved candidates
  • improve filtering before increasing top-k
  • reduce duplicated or low-value context
  • rerank fewer but better candidates
  • shorten context payloads before weakening grounding requirements
10. Define production acceptance criteria

A RAG system is ready for wider use only when it has:

  • a versioned eval set
  • retrieval metrics on representative query classes
  • a documented metadata schema
  • source traceability and citation behavior
  • ACL or tenant controls where needed
  • a known refresh or reindex policy
  • a troubleshooting path for common failures

Troubleshooting

Symptom: The answer sounds fluent but cites weak or irrelevant evidence

Check:

  • whether the expected passage appears in retrieved results at all
  • whether low-quality chunks outrank better ones
  • whether context packing includes too many marginal chunks
  • whether prompt instructions are causing overconfident synthesis

Likely fixes:

  • improve retrieval recall first
  • tighten chunk boundaries
  • add reranking
  • reduce noisy context
  • require the system to narrow claims when evidence is weak
Symptom: Exact identifiers are missed

Common causes:

  • semantic-only retrieval on identifier-heavy data
  • normalization that strips meaningful tokens
  • missing lexical path
  • poor metadata filtering

Likely fixes:

  • add keyword or hybrid retrieval
  • preserve exact identifiers in chunks and metadata
  • test identifier queries as their own eval class
Symptom: Relevant documents are found, but wrong sections are used

Common causes:

  • chunks too large or too mixed
  • sections not preserved during ingestion
  • reranker or context assembly preferring broad summaries

Likely fixes:

  • chunk on section boundaries
  • retain headings and local path metadata
  • rerank for passage relevance, not only document relevance
Symptom: Retrieval returns many near-duplicates

Common causes:

  • duplicated source documents
  • overlapping chunks dominating top results
  • repeated boilerplate text

Likely fixes:

  • deduplicate during ingestion
  • collapse near-duplicate neighbors in ranking
  • downweight boilerplate-heavy chunks
Symptom: Good retrieval offline, poor answers online

Common causes:

  • online filters differ from eval conditions
  • context truncation removes the best evidence
  • answer stage ignores or misuses evidence
  • stale index or stale metadata in production

Likely fixes:

  • compare offline and online traces
  • verify final packed context
  • check citation-to-source mapping
  • confirm refresh and reindex behavior
Symptom: Cross-tenant or restricted content leakage risk

Common causes:

  • filters applied after retrieval instead of before or during it
  • missing ACL metadata at chunk level
  • unsafe logs that expose retrieved text

Likely fixes:

  • enforce authorization in retrieval
  • carry ACL metadata into every chunk
  • sanitize traces and debug output
Symptom: The system is slow or too expensive

Common causes:

  • over-retrieval
  • expensive reranking depth
  • oversized context assembly
  • unnecessary second-pass calls

Likely fixes:

  • reduce candidate count with better filtering
  • rerank only where recall already looks acceptable
  • pass fewer, better chunks to the model
  • set explicit per-stage budgets

For a more detailed symptom-to-fix matrix, use references/domain-notes.md.

Examples

Open examples/worked-example.md for a concrete mini-corpus showing:

  • corpus preparation
  • metadata fields
  • document-aware chunking
  • retrieval eval queries
  • expected retrieval behavior
  • failure analysis
  • before/after tuning decisions

Additional Resources

Consider a different or additional skill when the center of gravity changes:

  • use a database or search-infrastructure skill when the main work is storage engine administration or production database operations
  • use an eval-focused skill when the main task is dataset design, scoring, and regression automation across many systems
  • use an application security skill when the main issue is data isolation, authorization architecture, or compliance review beyond retrieval boundaries

Execution Notes

During execution, keep outputs concrete:

  • name the corpus scope
  • list the metadata fields
  • describe the retrieval path
  • identify the eval queries used
  • state what changed and why
  • distinguish retrieval fixes from answer-generation fixes

A strong final answer from this skill should leave the operator with a retrieval plan that can be tested, traced, and improved without guesswork.

© diegosouzapw, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 8 other files (scripts, references) in skills_omni/rag-engineer of diegosouzapw/awesome-omni-skills.

  • SKILL.md
  • ATTRIBUTION.md
  • OMNI_ENHANCED.json
  • agents/openai.yaml
  • examples/worked-example.md
  • metadata.json
  • references/checklist.md
  • references/domain-notes.md
  • scripts/render_rag_review.py

Open the folder on GitHubat commit c3af004

Compare with similar skills

RAG Engineer next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

RAG Engineer compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
RAG Engineer this skilldiegosouzapw/awesome-omni-skills159—~3.5kAutomated safety check: PassMIT
Chroma Vector DatabaseOrchestra-Research/AI-Research-SKILLs13k8 repos~2.3kAutomated safety check: PassMIT
LLM Application DevMoizIbnYousaf/ai-agent-skills1.1k2 repos~1.3kAutomated safety check: PassMIT
Senior Prompt Engineermaslennikov-ig/claude-code-orchestrator-kit2604 repos~1.4kAutomated safety check: PassCustom licence
MCP Local RAGshinpr/mcp-local-rag407—~4.4kAutomated safety check: PassMIT
Ms Agent Framework RAGshuyu-labs/WebCode278—~1.1kAutomated safety check: PassCustom licence

Similar skills

  • Chroma Vector Database

    Orchestra-Research/AI-Research-SKILLs

    Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.

    13k GitHub starsUsed in 8 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed
  • LLM Application Dev

    MoizIbnYousaf/ai-agent-skills

    Building applications with Large Language Models - prompt engineering, RAG patterns, and LLM integration.

    1.1k GitHub starsUsed in 2 repos~1.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Senior Prompt Engineer

    maslennikov-ig/claude-code-orchestrator-kit

    Provides reference guides and Python scripts for prompt optimization, RAG evaluation, and agent orchestration when building or tuning LLM systems.

    260 GitHub starsUsed in 4 repos~1.4k tokens
    AI & LLM EngineeringAuto-check passed
  • MCP Local RAG

    shinpr/mcp-local-rag

    Searches, saves, and maintains a local document index through a local RAG MCP server.

    407 GitHub stars~4.4k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Ms Agent Framework RAG

    shuyu-labs/WebCode

    Comprehensive guide for building Agentic RAG systems using Microsoft Agent Framework in C.

    278 GitHub stars~1.1k tokensUpdated 3 mo ago
    AI & LLM EngineeringAuto-check passed
  • Local RAG Search

    nkapila6/mcp-local-rag

    Efficiently perform web searches using the mcp-local-rag server with semantic similarity ranking.

    134 GitHub starsUsed in 1 repo~1.6k tokens
    AI & LLM EngineeringAuto-check passed

More from diegosouzapw/awesome-omni-skills

All 39 skills in this repo
  • Content Creator

    diegosouzapw/awesome-omni-skills

    Content Creator workflow skill. An agent skill from diegosouzapw/awesome-omni-skills.

    159 GitHub stars~4k tokensUpdated 3 mo ago
    Auto-check passed
  • Helm Chart Scaffolding

    diegosouzapw/awesome-omni-skills

    Helm Chart Scaffolding workflow skill. An agent skill from diegosouzapw/awesome-omni-skills.

    159 GitHub stars~2.4k tokensUpdated 3 mo ago
    Auto-check passed
  • Prompt Engineering

    diegosouzapw/awesome-omni-skills

    Prompt Engineering Patterns workflow skill. An agent skill from diegosouzapw/awesome-omni-skills.

    159 GitHub stars~3.4k tokensUpdated 3 mo ago
    Auto-check passed
  • Prompt Engineering Patterns

    diegosouzapw/awesome-omni-skills

    Prompt Engineering Patterns workflow skill. An agent skill from diegosouzapw/awesome-omni-skills.

    159 GitHub stars~4k tokensUpdated 3 mo ago
    Auto-check passed
  • Prompt Library

    diegosouzapw/awesome-omni-skills

    📝 Prompt Library workflow skill. An agent skill from diegosouzapw/awesome-omni-skills.

    159 GitHub stars~3.1k tokensUpdated 3 mo ago
    Auto-check passed
  • Protocol Reverse Engineering

    diegosouzapw/awesome-omni-skills

    Protocol Reverse Engineering workflow skill. An agent skill from diegosouzapw/awesome-omni-skills.

    159 GitHub stars~3.8k tokensUpdated 3 mo ago
    Auto-check passed

Questions about RAG Engineer

What does RAG Engineer do?

RAG workflow skill. An agent skill from diegosouzapw/awesome-omni-skills. RAG Engineer is an agent skill from diegosouzapw/awesome-omni-skills. RAG workflow skill.

When should I use RAG Engineer?

RAG Engineer fits situations like: A user needs retrieval pipelines; evaluation for an AI application.

How do I install RAG Engineer in Claude Code?

Run `npx skills add diegosouzapw/awesome-omni-skills --skill rag-engineer -a claude-code`. Or copy the skill folder (skills_omni/rag-engineer in diegosouzapw/awesome-omni-skills) into .claude/skills/rag-engineer in your project. Claude Code loads it when a task matches its description.

How do I install RAG Engineer in Codex?

Run `npx skills add diegosouzapw/awesome-omni-skills --skill rag-engineer -a codex`. Or copy the skill folder (skills_omni/rag-engineer in diegosouzapw/awesome-omni-skills) into .agents/skills/rag-engineer in your project. Codex loads it when a task matches its description.

Can I use RAG Engineer in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add diegosouzapw/awesome-omni-skills --skill rag-engineer -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/rag-engineer, .gemini/skills/rag-engineer, .github/skills/rag-engineer and .opencode/skills/rag-engineer in your project.

What does RAG Engineer need to run?

Going by SKILL.md and its folder, RAG Engineer needs Python for the scripts in its folder. Our summary lists: Python 3.

Does RAG Engineer access the network?

SKILL.md names 3 domains. As links in the text: platform.openai.com, developers.openai.com and github.com. This is read from the text; nothing was executed.

Is RAG Engineer safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does RAG Engineer use?

RAG Engineer is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does RAG Engineer use?

About 3.5k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.8k tokens, read only when the agent opens those files.

What are the alternatives to RAG Engineer?

Skills that share tags, products or a category with RAG Engineer: Chroma Vector Database (Orchestra-Research/AI-Research-SKILLs, 13k stars), LLM Application Dev (MoizIbnYousaf/ai-agent-skills, 1.1k stars), Senior Prompt Engineer (maslennikov-ig/claude-code-orchestrator-kit, 260 stars) and MCP Local RAG (shinpr/mcp-local-rag, 407 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains RAG Engineer?

diegosouzapw (a GitHub user) maintains it in diegosouzapw/awesome-omni-skills, which has 159 GitHub stars. The repository holds 39 skills in this directory. The repository was last updated on July 8, 2026.

Source: diegosouzapw/awesome-omni-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.