Agent skill

RAG Architect

by alirezarezvani in alirezarezvani/claude-skills

A skill your agent uses when the user asks to design a RAG pipeline, choose a chunking strategy or embedding model, pick a vector database, or evaluate retrieval quality (precision@k, recall@k, NDCG).

MITAuto-check passedAI & LLM Engineering

Install RAG Architect

skills CLI
$ npx skills add alirezarezvani/claude-skills --skill rag-architect -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install alirezarezvani/claude-skills rag-architect --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/alirezarezvani/claude-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/engineering/skills/rag-architect .claude/skills/rag-architect && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
rag-architect
GitHub stars
28k
Token cost
~1.1k tokens
SKILL.md length
424 words
Files
7 (incl. references)
Skills in repo
342
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when the user asks to design a RAG pipeline, choose a chunking strategy or embedding model, pick a vector database, or evaluate retrieval quality (precision@k, recall@k, NDCG).

  • Works in 4 steps: Analyze the corpus and pick chunking → Design the pipeline from requirements → Evaluate retrieval quality → …
  • The user asks to design a RAG pipeline
  • SKILL.md covers Hard rules, Embedding model tiers…, Workflow and References
  • Runs Python scripts from its folder; calls python3

What it does

RAG Architect is an agent skill from alirezarezvani/claude-skills. Use when the user asks to design a RAG pipeline, choose a chunking strategy or embedding model, pick a vector database, or evaluate retrieval quality (precision@k, recall@k, NDCG). Examples: 'design a RAG system for our docs', 'what chunk size should I use for this corpus', 'evaluate my retriever against ground truth'. NOT for general LLM cost tuning (use llm-cost-optimizer) or agent loops over retrieval (use agenthub).

Its SKILL.md is about 1.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files, including reference files (for example `chunking_optimizer.py`, `rag_pipeline_designer.py` and `references/chunking_strategies_comparison.md`).

It sits in AI & LLM Engineering, covering Retrieval-augmented generation, Embeddings and Vector databases. The repository describes itself as: 380 Claude Code skills & agent skills & plugins (30+ Agents, 70+ custom commands, 380+ skills, customizable references, scripts)for Claude Code, Codex, Gemini CLI, Cursor, and 8… The licence is MIT.

When your agent uses it

  • The user asks to design a RAG pipeline
  • Choose a chunking strategy
  • Embedding model
  • Pick a vector database

Example prompts

  • “design a RAG system for our docs”
  • “what chunk size should I use for this corpus”
  • “evaluate my retriever against ground truth”
  • “/rag-architect”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Analyze the corpus and pick chunking
  2. Design the pipeline from requirements
  3. Evaluate retrieval quality
  4. Verification loop

What it can do on your machine

Read from SKILL.md and the folder at commit 19392f7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

RAG Architect loads about 1.1k tokens when it runs, and up to ~11k if it reads all its reference files. Until then it costs about 109 tokens; SKILL.md has 424 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~109
When it runs · the whole SKILL.md, loaded when a task matches
~1.1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~11k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from alirezarezvani/claude-skills at commit 19392f7, republished under its MIT licence (© alirezarezvani). 424 words, ~1,121 tokens.

Download SKILL.mdSave it as .claude/skills/rag-architect/SKILL.md (or your agent's skills folder). This skill also uses 6 other files; get the full folder from GitHub.
name
rag-architect
description
Use when the user asks to design a RAG pipeline, choose a chunking strategy or embedding model, pick a vector database, or evaluate retrieval quality (precision@k, recall@k, NDCG). Examples: 'design a RAG system for our docs', 'what chunk size should I use for this corpus', 'evaluate my retriever against ground truth'. NOT for general LLM cost tuning (use llm-cost-optimizer) or agent loops over retrieval (use agenthub).

RAG Architect

Design, tune, and evaluate production RAG pipelines with three deterministic tools. Run the tools against the actual corpus and requirements — do not pick chunk sizes or databases by intuition.

Hard rules

  1. Never present model names or vendor prices as current facts. Embedding models and vector-DB pricing rot in months. Recommend a tier (see table below), name a current-generation candidate, and tell the user to verify against the provider's live pricing page.
  2. Every design ends with an evaluation run. A RAG design without retrieval_evaluator.py numbers is a hypothesis, not a deliverable.
  3. Chunking is corpus-driven. Run chunking_optimizer.py on the real documents before choosing a strategy.

Embedding model tiers (pattern, not price list)

TierCurrent-generation examples (verify before use)When
Fast / self-hostedall-MiniLM-L6-v2, bge-smallCost-sensitive, small scale, real-time
Balanced openall-mpnet-base-v2, bge-large, e5-largeQuality without API dependency
Quality APItext-embedding-3-large, voyage-3-largeAccuracy-priority general retrieval
Codevoyage-code-3, CodeBERT-familyCode search corpora

Pricing discipline: build the cost model with a placeholder table — columns model | $/1M tokens (verify) | dims | as-of date — and have the user fill in live numbers. Same for vector DBs (Pinecone/Weaviate/Qdrant/Chroma/pgvector): the selection criteria (managed vs self-hosted, scale, filtering, existing Postgres) are durable; the dollar figures are not.

Workflow

All paths relative to this skill folder. Outputs chain: corpus analysis → design → evaluation.

1. Analyze the corpus and pick chunking
bash
python3 chunking_optimizer.py /path/to/docs --extensions .md .txt -o chunking.json

Emits chunking.json with corpus_info, per-strategy strategy_results, a recommendation, and sample_chunks. Use recommendation.strategy and its config; show the user 2-3 sample_chunks so they can sanity-check boundaries.

Show full SKILL.md (178 more words)Show less
2. Design the pipeline from requirements

Write a requirements JSON with these keys (all required): document_types[], document_count, avg_document_size (chars), queries_per_day, query_patterns[], latency_requirement, budget_monthly, accuracy_priority (0-1), cost_priority (0-1), maintenance_complexity.

bash
python3 rag_pipeline_designer.py requirements.json -o design.json

Emits design.json with chunking, embedding, vector_db, retrieval, reranking, evaluation, total_cost, architecture_diagram (mermaid), and config_templates. Present the diagram; label every cost_monthly figure as an estimate to verify (rule 1).

3. Evaluate retrieval quality

Prepare queries.json (list of {id, text} or {"queries": [...]}) and ground_truth.json ({query_id: [relevant_doc_ids]}), then:

bash
python3 retrieval_evaluator.py queries.json /path/to/docs ground_truth.json --k-values 3 5 10 -o eval.json

Reports precision@k, recall@k, MRR, NDCG@k, plus poor_precision_examples / poor_recall_examples for failure analysis.

4. Verification loop

The design is done only when:

  1. eval.json meets targets — typical floors: precision@5 ≥ 0.8, recall@10 ≥ 0.85 (set per use case with the user).
  2. If below target: inspect the poor-example lists, then change one variable (chunking strategy → re-run step 1; embedding tier; add reranking; hybrid retrieval) and re-run step 3. Repeat.
  3. Every recommended model/price in the deliverable carries a "verify current pricing/model availability" note with an as-of date.

References

  • references/chunking_strategies_comparison.md — strategy trade-offs the optimizer implements
  • references/embedding_model_benchmark.md — benchmark methodology (dated snapshot; staleness warning at top)
  • references/rag_evaluation_framework.md — metric definitions (faithfulness, relevance, precision/recall/NDCG)

© alirezarezvani, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 6 other files (references) in engineering/skills/rag-architect of alirezarezvani/claude-skills.

  • SKILL.md
  • chunking_optimizer.py
  • rag_pipeline_designer.py
  • references/chunking_strategies_comparison.md
  • references/embedding_model_benchmark.md
  • references/rag_evaluation_framework.md
  • retrieval_evaluator.py

Open the folder on GitHubat commit 19392f7

Compare with similar skills

RAG Architect next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

RAG Architect compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
RAG Architect this skillalirezarezvani/claude-skills28k—~1.1kAutomated safety check: PassMIT
Chroma Vector DatabaseOrchestra-Research/AI-Research-SKILLs13k7 repos~2.3kAutomated safety check: PassMIT
Ms Agent Framework RAGshuyu-labs/WebCode278—~1.1kAutomated safety check: PassCustom licence
Pgvector Semantic Searchtimescale/pg-aiguide1.9k—~3.8kAutomated safety check: PassApache-2.0
RAG Implementationwshobson/agents40k9 repos~1.1kAutomated safety check: PassMIT
RAG Engineerdavila7/claude-code-templates32k5 repos~729Automated safety check: PassMIT

Similar skills

  • Chroma Vector Database

    Orchestra-Research/AI-Research-SKILLs

    Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.

    13k GitHub starsUsed in 7 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Ms Agent Framework RAG

    shuyu-labs/WebCode

    Comprehensive guide for building Agentic RAG systems using Microsoft Agent Framework in C.

    278 GitHub stars~1.1k tokensUpdated 3 mo ago
    AI & LLM EngineeringAuto-check passed
  • Pgvector Semantic Search

    timescale/pg-aiguide

    A skill your agent uses for setting up vector similarity search with pgvector for AI/ML embeddings, RAG applications, or semantic search.

    1.9k GitHub stars~3.8k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • RAG Implementation

    wshobson/agents

    Build retrieval-augmented generation systems: pick a vector database and embedding model, choose retrieval and reranking strategies, and start from a LangGraph pipeline.

    40k GitHub starsUsed in 9 repos~1.1k tokens
    AI & LLM EngineeringAuto-check passed
  • RAG Engineer

    davila7/claude-code-templates

    Expert in building Retrieval-Augmented Generation systems. An agent skill from davila7/claude-code-templates.

    32k GitHub starsUsed in 5 repos~729 tokens
    AI & LLM EngineeringAuto-check passed
  • RAG Architect

    Jeffallan/claude-skills

    Designs retrieval-augmented generation systems: document chunking, embeddings, vector store setup, hybrid search, reranking and retrieval evaluation, with checks at each step.

    12k GitHub stars~2k tokensUpdated 5 days ago
    AI & LLM EngineeringAuto-check passed

More from alirezarezvani/claude-skills

All 342 skills in this repo
  • Agile Product Owner

    alirezarezvani/claude-skills

    Writes INVEST-checked user stories with acceptance criteria, splits epics, plans sprints from velocity and ranks the backlog with a weighted score.

    28k GitHub starsUsed in 3 repos~3.2k tokens
    Auto-check passed
  • Product Strategist

    alirezarezvani/claude-skills

    OKR cascade toolkit for product leaders: generates aligned company-to-team OKRs from five strategy types and scores how well they line up.

    28k GitHub starsUsed in 2 repos~1.8k tokens
    Auto-check passed
  • App Store Optimization

    alirezarezvani/claude-skills

    App Store Optimization (ASO) toolkit for researching keywords, analyzing competitor rankings, generating metadata suggestions, and improving app visibility on Apple App Store and Google Play Store.

    28k GitHub starsUsed in 1 repo~4.2k tokens
    Auto-check passed
  • AWS Solution Architect

    alirezarezvani/claude-skills

    Design AWS architectures for startups using serverless patterns and IaC templates.

    28k GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • Campaign Analytics

    alirezarezvani/claude-skills

    Calculates attribution, funnel and ROI figures for marketing campaigns with three Python scripts that need only the standard library.

    28k GitHub starsUsed in 1 repo~2.1k tokens
    Auto-check passed
  • Code to PRD

    alirezarezvani/claude-skills

    Reverse-engineers a frontend, backend or fullstack codebase into a product requirements document with per-page docs, an enum dictionary and an API inventory.

    28k GitHub starsUsed in 1 repo~4.9k tokens
    Auto-check passed

Questions about RAG Architect

What does RAG Architect do?

A skill your agent uses when the user asks to design a RAG pipeline, choose a chunking strategy or embedding model, pick a vector database, or evaluate retrieval quality (precision@k, recall@k, NDCG). RAG Architect is an agent skill from alirezarezvani/claude-skills. Use when the user asks to design a RAG pipeline, choose a chunking strategy or embedding model, pick a vector database, or evaluate retrieval quality (precision@k, recall@k, NDCG).

When should I use RAG Architect?

RAG Architect fits situations like: the user asks to design a RAG pipeline; choose a chunking strategy; embedding model; pick a vector database.

How do I install RAG Architect in Claude Code?

Run `npx skills add alirezarezvani/claude-skills --skill rag-architect -a claude-code`. Or copy the skill folder (engineering/skills/rag-architect in alirezarezvani/claude-skills) into .claude/skills/rag-architect in your project. Claude Code loads it when a task matches its description.

How do I install RAG Architect in Codex?

Run `npx skills add alirezarezvani/claude-skills --skill rag-architect -a codex`. Or copy the skill folder (engineering/skills/rag-architect in alirezarezvani/claude-skills) into .agents/skills/rag-architect in your project. Codex loads it when a task matches its description.

Can I use RAG Architect in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add alirezarezvani/claude-skills --skill rag-architect -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/rag-architect, .gemini/skills/rag-architect, .github/skills/rag-architect and .opencode/skills/rag-architect in your project.

What does RAG Architect need to run?

Going by SKILL.md and its folder, RAG Architect needs Python for the scripts in its folder and the command-line tools its instructions call (python3). Our summary lists: Python 3.

Does RAG Architect access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is RAG Architect safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does RAG Architect use?

RAG Architect is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does RAG Architect use?

About 1.1k tokens (SKILL.md is roughly 4.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 9.5k tokens, read only when the agent opens those files.

What are the alternatives to RAG Architect?

Skills that share tags, products or a category with RAG Architect: Chroma Vector Database (Orchestra-Research/AI-Research-SKILLs, 13k stars), Ms Agent Framework RAG (shuyu-labs/WebCode, 278 stars), Pgvector Semantic Search (timescale/pg-aiguide, 1.9k stars) and RAG Implementation (wshobson/agents, 40k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains RAG Architect?

alirezarezvani (a GitHub user) maintains it in alirezarezvani/claude-skills, which has 27,891 GitHub stars. The repository holds 342 skills in this directory. The repository was last updated on August 30, 2026.

Source: alirezarezvani/claude-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.