Agent skill

RAG Accuracy Optimizer

by LeoYeAI in LeoYeAI/openclaw-master-skills

Optimize accuracy for RAG (Retrieval-Augmented Generation) systems.

MITAuto-check passedAI & LLM Engineering

Install RAG Accuracy Optimizer

skills CLI
$ npx skills add LeoYeAI/openclaw-master-skills --skill rag-accuracy-optimizer -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install LeoYeAI/openclaw-master-skills rag-accuracy-optimizer --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/LeoYeAI/openclaw-master-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/rag-accuracy-optimizer .claude/skills/rag-accuracy-optimizer && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
rag-accuracy-optimizer
GitHub stars
2.2k
Token cost
~5.5k tokens
SKILL.md length
1,520 words
Files
14 (incl. scripts, references)
Skills in repo
972
Repo updated
First seen
Licence
MIT

At a glance

Optimize accuracy for RAG (Retrieval-Augmented Generation) systems.

  • Works in 11 steps: Structured Data Design → Chunking Strategies → Retrieval Optimization → …
  • Improving a RAG pipeline
  • SKILL.md covers Workflow Overview, 1. Structured Data Design, 2. Chunking Strategies and 3. Retrieval Optimization, plus 2 more sections
  • Runs Python scripts from its folder; calls python3, pip and just

What it does

RAG Accuracy Optimizer is an agent skill from LeoYeAI/openclaw-master-skills. Optimize accuracy for RAG (Retrieval-Augmented Generation) systems. Covers: DB schema design, chunking strategies, retrieval optimization, accuracy testing, and anti-hallucination safeguards. Use when: (1) designing or improving a RAG pipeline, (2) choosing the right chunking strategy, (3) optimizing retrieval accuracy (hybrid search, reranking, multi-query), (4) evaluating chunk quality or testing accuracy, (5) setting up monitoring & safeguards for RAG production, (6) choosing SQL vs Vector DB, (7) designing…

Its SKILL.md is about 5.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 15 other files, including scripts and reference files (for example `_meta.json`, `references/advanced-rag.md` and `references/chunking-patterns.md`).

It sits in AI & LLM Engineering, covering Retrieval-augmented generation. It works with SQL. The repository describes itself as: 🧠 Curated collection of 1209+ best OpenClaw skills — weekly updated by MyClaw.ai. The licence is MIT.

When your agent uses it

  • Improving a RAG pipeline
  • Choosing the right chunking strategy
  • Optimizing retrieval accuracy (hybrid search
  • Evaluating chunk quality

Example prompts

  • “/rag-accuracy-optimizer”

Requirements

  • Python 3

Workflow steps

11 steps, taken from the step headings in SKILL.md.

  1. Structured Data Design
  2. Chunking Strategies
  3. Retrieval Optimization
  4. Accuracy Testing & Monitoring
  5. Safeguards
  6. Embedding Model Selection
  7. Vector DB Comparison
  8. Advanced Techniques
  9. Performance Optimization
  10. Vietnamese-Specific RAG
  11. AI Orchestrator — Multi-Model Cost Optimization

What it can do on your machine

Read from SKILL.md and the folder at commit e5199b5. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 4 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3
    • pip
    • just

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

RAG Accuracy Optimizer loads about 5.5k tokens when it runs, and up to ~34k if it reads all its reference files. Until then it costs about 157 tokens; SKILL.md has 1,520 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~157
When it runs · the whole SKILL.md, loaded when a task matches
~5.5k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~34k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from LeoYeAI/openclaw-master-skills at commit e5199b5, republished under its MIT licence (© LeoYeAI). 1,520 words, ~5,456 tokens.

Download SKILL.mdSave it as .claude/skills/rag-accuracy-optimizer/SKILL.md (or your agent's skills folder). This skill also uses 13 other files; get the full folder from GitHub.
name
rag-accuracy-optimizer
description
Optimize accuracy for RAG (Retrieval-Augmented Generation) systems. Covers: DB schema design, chunking strategies, retrieval optimization, accuracy testing, and anti-hallucination safeguards. Use when: (1) designing or improving a RAG pipeline, (2) choosing the right chunking strategy, (3) optimizing retrieval accuracy (hybrid search, reranking, multi-query), (4) evaluating chunk quality or testing accuracy, (5) setting up monitoring & safeguards for RAG production, (6) choosing SQL vs Vector DB, (7) designing metadata schemas for domain-specific data (insurance, finance, healthcare, e-commerce).

RAG Accuracy Optimizer

A skill for optimizing end-to-end accuracy in RAG systems.

Workflow Overview

Data Design → Chunking → Indexing → Retrieval → Generation → Testing → Monitoring

Each step impacts accuracy. Optimize each step in order.


1. Structured Data Design

SQL vs Vector DB — When to Use What?
CriteriaSQL (PostgreSQL, MySQL)Vector DB (Pinecone, Qdrant, Weaviate)
Exact facts (price, date, product code)✅ Optimal❌ Not suitable
Semantic search (query meaning)❌ Not supported✅ Optimal
Aggregation (SUM, COUNT, AVG)✅ Native❌ Not supported
Fuzzy matching ("similar to...")⚠️ Limited✅ Optimal
Hybrid (recommended)pgvector for bothVector DB + SQL metadata store

Principle: Clearly structured data → SQL. Unstructured data requiring semantic understanding → Vector DB. Most production systems need both.

Schema Design Patterns by Domain

Insurance:

policies(policy_id, product_type, effective_date)
clauses(clause_id, policy_id, clause_number, title, content)
exclusions(exclusion_id, clause_id, description)
-- Vector: embedding for clause.content + exclusion.description

Finance:

securities(ticker, name, sector, exchange)
reports(report_id, ticker, period, report_type)
sections(section_id, report_id, heading, content)
-- Vector: embedding for section.content, metadata: ticker + period

Healthcare:

drugs(drug_id, generic_name, brand_name, category)
guidelines(guideline_id, condition, recommendation, evidence_level)
interactions(drug_a_id, drug_b_id, severity, description)
-- Vector: embedding for guidelines.recommendation

E-commerce:

products(product_id, name, category, brand, price)
reviews(review_id, product_id, rating, content)
specs(product_id, attribute, value)
-- Vector: embedding for review.content + product description
Metadata Tagging Strategy

Each chunk/document needs at minimum:

python
metadata = {
    "source": "policy_doc_v2.pdf",       # Origin
    "source_type": "pdf",                 # File type
    "domain": "insurance",                # Domain
    "category": "life_insurance",          # Classification
    "entity_id": "POL-2024-001",          # Related entity ID
    "section": "exclusions",              # Section in doc
    "chunk_index": 3,                      # Chunk position
    "total_chunks": 12,                    # Total chunks in doc
    "created_at": "2024-01-15",           # Creation date
    "version": "2.0",                      # Version
    "language": "en"                       # Language
}

Metadata principles:

  • Always include source for traceability and citation
  • entity_id enables pre-filtering before search → reduces noise
  • chunk_index + total_chunks enables fetching surrounding context
  • Domain-specific fields (clause_number, ticker, drug_id) vary by use case
Normalization vs Denormalization
NormalizedDenormalized
ProsLess duplication, easy to updateFaster queries, fewer JOINs
ConsRequires JOINs, slowerDuplication, harder to sync
Use whenSource of truth (SQL)Vector store chunks

Recommendation: Normalized for SQL source → Denormalized when creating chunks for Vector DB. Each chunk should contain sufficient context, no JOINs needed at retrieval time.


2. Chunking Strategies

Detailed code examples: read references/chunking-patterns.md

Choosing the Right Strategy
Data has clear structure (clauses, sections)?
  → Semantic chunking (by heading/section)

Long, continuous data (articles, transcripts)?
  → Fixed size + overlap (512 tokens, 10-20% overlap)

Need both overview + detail?
  → Hierarchical chunking (parent-child)

Domain-specific with its own logical units?
  → Domain-specific chunking
Chunk Size Guidelines
SizeUse caseTrade-off
128-256 tokensFAQ, short definitionsHigh precision, less context
256-512 tokensRecommended defaultGood balance
512-1024 tokensComplex text, legal docsMore context, potential noise
>1024 tokensRarely usedToo much noise
Semantic Chunking

Split by meaning (section, topic) instead of fixed size:

python
# Split by markdown headings
# Split by paragraph breaks (\n\n)
# Split by topic change (using NLP or LLM detection)
Overlap Strategy
  • 10-20% overlap between adjacent chunks
  • Ensures information at boundaries is not lost
  • Chunk N ends with 1-2 opening sentences of chunk N+1
Hierarchical Chunking (Parent-Child)
Document (summary)
  └── Section (heading + key points)
        └── Paragraph (details)
  • Search at paragraph level (most detailed)
  • When matched, pull parent section for additional context
  • Keep parent_id in metadata
Domain-Specific Chunking
  • Insurance: 1 chunk = 1 clause
  • Finance: 1 chunk = 1 report section, metadata = ticker + period
  • Healthcare: 1 chunk = 1 guideline/recommendation
  • E-commerce: 1 chunk = 1 review or 1 product description
  • Legal: 1 chunk = 1 article/clause/section
Metadata Enrichment Per Chunk

Each chunk should be enriched with:

  • Summary: 1-2 sentence content summary (LLM-generated)
  • Keywords: Key terms (supports BM25)
  • Questions: 2-3 questions this chunk can answer (hypothetical questions)
  • Entities: Named entities (product names, codes, dates)

3. Retrieval Optimization

Detailed code examples: read references/retrieval-patterns.md

User Query
  → Query Rewriting (expand/reformulate)
  → Multi-Query Generation (3-5 variants)
  → Metadata Filtering (narrow scope)
  → Hybrid Search (Vector + BM25)
  → Merge & Deduplicate
  → Reranking (top 20 → top 5)
  → Contextual Compression
  → LLM Generation (with citations)
Hybrid Search (Vector + BM25)
  • Vector search: Find by meaning (semantic similarity)
  • BM25 (keyword): Find by exact keywords (product names, codes)
  • Combined: Weighted fusion or Reciprocal Rank Fusion (RRF)
final_score = α × vector_score + (1-α) × bm25_score
# α = 0.7 is a good starting point, tune per domain
Query Rewriting

Use LLM to reformulate the user question for clarity:

User: "does insurance pay?"
→ Rewritten: "Under what circumstances does life insurance pay out benefits?"
Multi-Query

From 1 question, generate 3-5 variants → search each variant → merge results:

Original: "Which bank has the highest savings rate?"
Query 1: "Compare savings interest rates across banks 2024"
Query 2: "Bank with highest deposit rate currently"
Query 3: "Top banks with best deposit interest rates"
Reranking

After retrieval, use a reranking model to re-sort by relevance:

  • Cohere Rerank: Simple API, highly effective
  • Cross-encoder: More accurate than bi-encoder, but slower
  • GPT Rerank: Use LLM to evaluate relevance (expensive but flexible)

Retrieve top 20 → rerank → take top 3-5 for generation.

Contextual Compression

After reranking, compress each chunk: keep only the part relevant to the question.

Original chunk (500 tokens) → Compressed (150 tokens, relevant part only)

Reduces noise, saves context window, improves accuracy.

Metadata Filtering

Narrow the search space BEFORE vector search:

python
# Instead of searching all 1M chunks:
filter = {"domain": "insurance", "product_type": "life"}
# Only search within ~50K relevant chunks
results = vector_db.search(query, filter=filter, top_k=20)

4. Accuracy Testing & Monitoring

Test Suite Design

Create ground truth Q&A pairs:

json
{
    "test_cases": [
        {
            "question": "Does life insurance pay out for suicide?",
            "expected_answer": "No payout within the first 2 years",
            "expected_source": "clause_15_exclusions.pdf",
            "category": "exclusions",
            "difficulty": "medium"
        }
    ]
}

Recommendation: Minimum 50-100 test cases, evenly distributed across categories and difficulty levels.

Metrics
MetricMeaningTarget
Precision@K% relevant results in top K>0.8
Recall@K% ground truth found in top K>0.9
F1Harmonic mean of Precision and Recall>0.85
MRRMean Reciprocal Rank — average position of first correct result>0.8
NDCGNormalized Discounted Cumulative Gain — ranking quality>0.85
Answer Accuracy% correct answers (human eval or LLM judge)>0.9
A/B Testing

Compare strategies by running the same test suite:

Config A: chunk_size=256, overlap=10%, no_rerank
Config B: chunk_size=512, overlap=20%, cohere_rerank
→ Compare MRR, NDCG, Answer Accuracy
→ Choose the config with better metrics
Error Analysis Framework

Classify errors to know where to optimize:

Error TypeCauseSolution
Retrieval MissCorrect chunk not foundImprove chunking, add hypothetical Q
Ranking ErrorCorrect chunk found but ranked lowAdd reranking
Generation ErrorCorrect chunk but LLM answers wrongImprove prompt, add few-shot
No AnswerInformation not in DBExpand knowledge base
HallucinationLLM fabricates informationAdd citation enforcement
Production Monitoring

Log each query:

python
log_entry = {
    "timestamp": "2024-01-15T10:30:00",
    "query": "...",
    "retrieved_chunks": [...],
    "reranked_chunks": [...],
    "answer": "...",
    "confidence": 0.85,
    "latency_ms": 450,
    "user_feedback": None  # thumbs up/down
}

Alerts:

  • Continuous confidence < 0.5 → review chunking/retrieval
  • Latency > 2s → optimize index or reduce top_k
  • Negative feedback > 20% → audit error patterns

5. Safeguards

Hallucination Prevention

Mandatory system prompt:

Answer ONLY based on the information provided in the context.
If you cannot find the information, respond: "I could not find this
information in the available data."
NEVER fabricate information.
Citation Enforcement

Require source citations:

Every answer must include [Source: file_name, section/clause].
If a specific source cannot be cited, mark it as "unverified".
Confidence Thresholds
python
if max_relevance_score < 0.3:
    return "No relevant information found."
elif max_relevance_score < 0.6:
    return answer + "\n⚠️ Low confidence. Please verify."
else:
    return answer + f"\n📎 Source: {sources}"
Answer Verification

Cross-check the answer with the DB:

  1. Extract claims from the answer (using LLM)
  2. Verify each claim against retrieved chunks
  3. Flag claims without supporting evidence
  4. Return only verified claims

6. Embedding Model Selection

Detailed comparison: read references/embedding-models.md

Quick Decision
ScenarioModelReason
Production, budget OKCohere embed-v4Highest MTEB, input_type optimization
Production, low costOpenAI text-embedding-3-small$0.02/1M tokens, good quality
Self-host, multilingualBGE-M3 ⭐Hybrid dense+sparse, 100+ languages, free
Self-host, VietnameseBGE-M3 or multilingual-e5-largeBest for Vietnamese RAG
POC / Prototypeall-MiniLM-L6-v290MB, runs on CPU
Key Principles
  • Dimension reduction: OpenAI embed-3 supports Matryoshka — reduce 3072→512 with only ~3% quality loss
  • Normalize embeddings: Always normalize_embeddings=True when encoding for cosine similarity
  • Batch processing: Encode in batches (256-2000 items) instead of one at a time
  • Consistency: Use the SAME model for indexing and querying

7. Vector DB Comparison

Detailed comparison + HNSW tuning: read references/vector-db-comparison.md

Quick Decision
Already have PostgreSQL and <5M vectors? → pgvector
Just prototype/POC? → ChromaDB
Production, want zero-ops? → Pinecone
Need performance + HNSW control? → Qdrant
Need hybrid BM25+vector built-in? → Weaviate
HNSW Tuning Quick Reference
ParamDefaultAccuracy-criticalSpeed-critical
M1648-648-16
ef_construction200400-500100-200
ef (search)100200-25650-100

Trade-off: Higher M and ef → better recall but more RAM and slower. Tune per SLA.


8. Advanced Techniques

Detailed code examples: read references/advanced-rag.md

Show full SKILL.md (619 more words)Show less
Late Chunking

Embed the entire document first, then pool embeddings by chunk boundaries. Each chunk retains context from surrounding text.

Traditional: Doc → Chunk → Embed each (loses context)
Late Chunking: Doc → Embed full → Pool by boundaries (retains context)

Use when: Documents have many co-references ("it", "this", "the package"). Quality gain: +5-10%.

RAPTOR (Recursive Abstractive Processing)

Build a multi-level summary tree: Level 0 (chunks) → Level 1 (summaries) → Level 2 (summary of summaries).

Use when: Need to answer both broad queries ("Compare all insurance packages") and narrow queries ("Clause X of Package Y"). Quality gain: +10-15%.

GraphRAG (Microsoft)

Build a knowledge graph from documents → detect communities → summarize communities → query via map-reduce.

Use when: Multi-hop reasoning, synthesize across many documents. Quality gain: +15-25% for synthesis queries. High overhead (many LLM calls when building the graph).

Combining Techniques (Production Stack)
1. Late Chunking → better embeddings
2. Hybrid Search (BM25 + vector) → high recall
3. Reranking (Cohere/Cross-encoder) → high precision
4. RAPTOR → multi-level retrieval (optional)
5. GraphRAG → synthesis queries (optional, high cost)

9. Performance Optimization

Caching Layer
python
# Cache embeddings (avoid re-computation)
import hashlib, json, redis

r = redis.Redis()

def cached_embed(text, model):
    key = f"emb:{hashlib.md5(text.encode()).hexdigest()}"
    cached = r.get(key)
    if cached:
        return json.loads(cached)
    embedding = model.encode([text])[0].tolist()
    r.setex(key, 3600, json.dumps(embedding))  # TTL 1h
    return embedding

# Cache search results (avoid re-searching)
def cached_search(query, search_fn, ttl=300):
    key = f"search:{hashlib.md5(query.encode()).hexdigest()}"
    cached = r.get(key)
    if cached:
        return json.loads(cached)
    results = search_fn(query)
    r.setex(key, ttl, json.dumps(results))
    return results
Async Retrieval
python
import asyncio

async def parallel_retrieve(query, retrievers):
    """Run multiple retrievers in parallel."""
    tasks = [r.search(query) for r in retrievers]
    results = await asyncio.gather(*tasks)
    return merge_and_deduplicate(results)
HNSW Index Tuning

See details in references/vector-db-comparison.md HNSW section. Key: tune ef (search) per latency SLA, tune M per recall target.


10. Vietnamese-Specific RAG

Details: read references/vietnam-nlp.md

Key Challenges
IssueSolution
Diacritics (with vs without)Dual indexing: index both versions
Compound words ("bảo hiểm")Word segmentation (underthesea)
Abbreviations (BHXH, TTCK, BLLĐ)Abbreviation expansion dictionary
Vietnamese proper namesNER with underthesea/PhoBERT
Domain terms (finance, law, medical)Domain-specific term enrichment
Embedding Models for Vietnamese
  • BGE-M3: Best overall — hybrid dense+sparse, 100+ languages
  • multilingual-e5-large: Good alternative — retrieval-optimized
  • PhoBERT-v2: Best for NER/classification (needs fine-tuning for retrieval)
Preprocessing Pipeline
Input text
  → Unicode normalize (NFC)
  → Expand abbreviations (BHXH → Social Insurance)
  → Domain term enrichment
  → Dual index: original + no-diacritics version
  → Extract entities → metadata

11. AI Orchestrator — Multi-Model Cost Optimization

Detailed prompt templates, code examples: read references/orchestrator-patterns.md

Query Classification Pipeline

Each user query is classified into 1 of 5 categories:

CategoryDescriptionExampleModel
simpleGreeting, FAQ, simple lookup"Hello", "Opening hours?"No LLM / Local
ragNeeds knowledge base search"Does insurance cover cancer?"Cheap (Gemini Flash)
complexMulti-hop reasoning, comparison, analysis"Compare 3 insurance packages for a family of 4"Standard (GPT-4o-mini) / Premium (Claude Sonnet)
actionNeeds tool/API execution (create form, calculate)"Calculate insurance premium for me, age 30"Standard + Tools
unsafeViolation content, injection, jailbreak"Ignore instructions..."Block — No LLM
2-Stage Classification (Minimize LLM Tokens)
User Query
  → Stage 1: Rule-based pre-classifier (regex, keywords, NO LLM)
    → confidence ≥ 0.8? → DONE (skip LLM)
    → confidence < 0.8? → Stage 2: LLM classifier (cheap model, ~50 tokens)

Stage 1 blocks 60-80% of queries without spending a single LLM token.

Model Routing
Category → Model Selection:
  greeting/simple  → No LLM (rule-based response)
  rag (simple)     → Gemini Flash ($0.075/1M input) — cheap, fast
  rag (complex)    → GPT-4o-mini ($0.15/1M input) — balanced
  complex          → Claude Sonnet ($3/1M input) — premium quality
  action           → Gemini Flash + Tool calls
  unsafe           → Block response (no LLM cost)
Cost Optimization Rules
  1. Rule-based first: Greeting, FAQ, unsafe → DON'T call LLM
  2. Cheapest sufficient model: Prefer Gemini Flash for RAG queries
  3. Escalate on failure: Gemini Flash fail/low-confidence → GPT-4o-mini → Claude Sonnet
  4. Cache responses: Identical queries → cached answer (TTL 5-30 min)
  5. Batch classify: Multiple queries → 1 LLM call to classify all
  6. Token budget: Set max_tokens per category (simple: 100, rag: 300, complex: 500)
RAG Trigger Rules
ConditionRAG On/Off
Query contains domain keywords✅ ON
Classification = "rag" or "complex"✅ ON
Greeting, simple lookup, unsafe❌ OFF
Confidence score > 0.9 from cache/FAQ❌ OFF (answer from cache)
Tool Trigger Rules
ConditionTools
Query requests calculation (fees, interest)calculator tool
Query requests form creation/submissionform_builder tool
Query requests real-time lookup (price, exchange rate)api_lookup tool
Classification ≠ "action"No tools
JSON Output Format
json
{
  "category": "rag",
  "confidence": 0.92,
  "risk_level": "low",
  "model": "gemini-flash",
  "rag_enabled": true,
  "tools": [],
  "max_tokens": 300,
  "reasoning": "User asks about insurance benefits — needs knowledge base search"
}

Scripts

eval_ragas.py

RAGAS evaluation pipeline. Run:

bash
python3 scripts/eval_ragas.py --test-file eval_dataset.json --output results.json
python3 scripts/eval_ragas.py --test-file eval_dataset.json --metrics faithfulness,answer_relevancy

Input: JSON file with test cases (question, answer, contexts, ground_truth). Output: metrics report + threshold checks. Requires: pip install ragas langchain-openai datasets

embedding_benchmark.py

Benchmark embedding models on a Vietnamese dataset. Run:

bash
python3 scripts/embedding_benchmark.py --models bge-m3,multilingual-e5 --dataset vi_pairs.json
python3 scripts/embedding_benchmark.py --models all --quick  # Use built-in test pairs

Input: JSON file with query-positive-negative pairs. Output: accuracy + latency comparison. Requires: pip install sentence-transformers numpy torch

chunk_optimizer.py

Evaluate chunk quality. Run:

bash
python3 scripts/chunk_optimizer.py --input chunks.jsonl --output report.json

Input: JSONL file, each line is {"text": "...", "metadata": {...}}. Output: quality report with scores.

accuracy_test.py

Test framework for RAG accuracy. Run:

bash
python3 scripts/accuracy_test.py --test-file tests.json --results-dir ./results

Input: JSON file with test cases (question, expected_answer, expected_source). Output: metrics report.


References

  • references/chunking-patterns.md — Python code examples for chunking strategies
  • references/retrieval-patterns.md — Code examples for hybrid search, reranking, multi-query
  • references/embedding-models.md — Detailed embedding model comparison (OpenAI, Cohere, BGE-M3, PhoBERT...)
  • references/vector-db-comparison.md — Vector DB comparison + HNSW tuning guide
  • references/advanced-rag.md — Late Chunking, RAPTOR, GraphRAG with code examples
  • references/testing-frameworks.md — RAGAS, LLM-as-Judge, Adversarial testing
  • references/vietnam-nlp.md — Vietnamese NLP: diacritics, abbreviations, NER, domain terms
  • references/orchestrator-patterns.md — Multi-model orchestrator: prompt templates, rule-based pre-classifier, cost comparison, fallback chain, monitoring

© LeoYeAI, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 13 other files (scripts, references) in skills/rag-accuracy-optimizer of LeoYeAI/openclaw-master-skills.

  • SKILL.md
  • _meta.json
  • references/advanced-rag.md
  • references/chunking-patterns.md
  • references/embedding-models.md
  • references/orchestrator-patterns.md
  • references/retrieval-patterns.md
  • references/testing-frameworks.md
  • references/vector-db-comparison.md
  • references/vietnam-nlp.md
  • scripts/accuracy_test.py
  • scripts/chunk_optimizer.py
  • scripts/embedding_benchmark.py
  • scripts/eval_ragas.py

Open the folder on GitHubat commit e5199b5

Compare with similar skills

RAG Accuracy Optimizer next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

RAG Accuracy Optimizer compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
RAG Accuracy Optimizer this skillLeoYeAI/openclaw-master-skills2.2k—~5.5kAutomated safety check: PassMIT
Aliyun Opensearch Searchcinience/alicloud-skills397—~1.1kAutomated safety check: PassMIT
Tanyuan Searchinfometa/workbuddyskills344—~1.2kAutomated safety check: PassNone
Snowflake Cortex AIMindrally/skills268—~2.3kAutomated safety check: PassApache-2.0
Postgrestimescale/pg-aiguide1.9k—~941Automated safety check: PassApache-2.0
DBoracle/skills873—~1.4kAutomated safety check: PassUPL-1.0

Similar skills

  • Aliyun Opensearch Search

    cinience/alicloud-skills

    A skill your agent uses when working with OpenSearch vector search edition via the Python SDK (ha3engine) to push documents and run HA/SQL searches.

    397 GitHub stars~1.1k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Tanyuan Search

    infometa/workbuddyskills

    腾讯探元文博检索工具集(Agentic RAG)。封装两个 HTTP API 为 Node.js 脚本,由 Agent 依据问题特征选择工具并构造 query: - search-relics(文物/世界遗产数据库 NL→SQL):适合结构化事实的详情、列表、统计与排行查询 -…

    344 GitHub stars~1.2k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Snowflake Cortex AI

    Mindrally/skills

    Reference for Snowflake Cortex AI Functions (AICOMPLETE, AICLASSIFY, AIEXTRACT, AIFILTER, etc.) and Cortex Search for building RAG applications entirely inside Snowflake.

    268 GitHub stars~2.3k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Postgres

    timescale/pg-aiguide

    A skill your agent uses for any PostgreSQL database work — table design, indexing, data types, constraints, extensions (pgvector, PostGIS, TimescaleDB), search, and migrations.

    1.9k GitHub stars~941 tokensUpdated today
    DatabasesAuto-check passed
  • DB

    oracle/skills

    Official

    Oracle Database guidance for SQL, PL/SQL, SQLcl, ORDS, Oracle Vector SDK, administration, app development, performance, security, migrations, and agent-safe database workflows.

    873 GitHub stars~1.4k tokensUpdated 2 days ago
    DatabasesAuto-check passed
  • Enterprise AI

    oracle/skills

    Official

    Oracle Enterprise AI guidance for building, deploying, securing, estimating cost for, and integrating AI models, agents, RAG, Responses API workflows, custom or imported models, fine-tuning, model…

    873 GitHub stars~1.4k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed

More from LeoYeAI/openclaw-master-skills

All 972 skills in this repo
  • DevOps Pipeline Management

    LeoYeAI/openclaw-master-skills

    Manages pipelines on a DevOps quality and efficiency platform through its OpenAPI: list workspaces and templates, create, update, run and cancel pipelines, and read run records.

    2.2k GitHub stars~4.2k tokensUpdated 2 mo ago
    Auto-check: notes
  • Feishu Document Collaboration

    LeoYeAI/openclaw-master-skills

    Patches OpenClaw's Feishu extension so an edited document triggers an isolated agent session that reads the doc and replies inline, turning it into a live chat space.

    2.2k GitHub stars~2k tokensUpdated 2 mo ago
    Auto-check passed
  • Files Memory System

    LeoYeAI/openclaw-master-skills

    Multi-context memory management system for OpenClaw agents with group-isolated storage, global shared memory, workspace organization, and group-specific skills isolation.

    2.2k GitHub stars~3.8k tokensUpdated 2 mo ago
    Auto-check passed
  • GEO-Claw AI Visibility Agent

    LeoYeAI/openclaw-master-skills

    Runs a brand's AI-search visibility work end to end: diagnosing how AI platforms represent it, repositioning it, producing AI-optimized content and monitoring ongoing mentions.

    2.2k GitHub stars~4.7k tokensUpdated 2 mo ago
    Auto-check passed
  • Google Workspace CLI

    LeoYeAI/openclaw-master-skills

    Installs and authenticates the gws CLI, then automates Gmail, Drive, Sheets, Calendar, Docs, Chat and Tasks with ready-made recipes, persona bundles and security audits.

    2.2k GitHub stars~2.6k tokensUpdated 2 mo ago
    Auto-check: notes
  • HealthFit Health Advisors

    LeoYeAI/openclaw-master-skills

    Runs four advisor roles, a fitness coach, nutritionist, data analyst and TCM practitioner, to build a health profile and track workouts, diet and wellness over time.

    2.2k GitHub stars~4.4k tokensUpdated 2 mo ago
    Auto-check passed

Works with

Questions about RAG Accuracy Optimizer

What does RAG Accuracy Optimizer do?

Optimize accuracy for RAG (Retrieval-Augmented Generation) systems. RAG Accuracy Optimizer is an agent skill from LeoYeAI/openclaw-master-skills. Optimize accuracy for RAG (Retrieval-Augmented Generation) systems.

When should I use RAG Accuracy Optimizer?

RAG Accuracy Optimizer fits situations like: improving a RAG pipeline; choosing the right chunking strategy; optimizing retrieval accuracy (hybrid search; evaluating chunk quality.

How do I install RAG Accuracy Optimizer in Claude Code?

Run `npx skills add LeoYeAI/openclaw-master-skills --skill rag-accuracy-optimizer -a claude-code`. Or copy the skill folder (skills/rag-accuracy-optimizer in LeoYeAI/openclaw-master-skills) into .claude/skills/rag-accuracy-optimizer in your project. Claude Code loads it when a task matches its description.

How do I install RAG Accuracy Optimizer in Codex?

Run `npx skills add LeoYeAI/openclaw-master-skills --skill rag-accuracy-optimizer -a codex`. Or copy the skill folder (skills/rag-accuracy-optimizer in LeoYeAI/openclaw-master-skills) into .agents/skills/rag-accuracy-optimizer in your project. Codex loads it when a task matches its description.

Can I use RAG Accuracy Optimizer in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add LeoYeAI/openclaw-master-skills --skill rag-accuracy-optimizer -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/rag-accuracy-optimizer, .gemini/skills/rag-accuracy-optimizer, .github/skills/rag-accuracy-optimizer and .opencode/skills/rag-accuracy-optimizer in your project.

What does RAG Accuracy Optimizer need to run?

Going by SKILL.md and its folder, RAG Accuracy Optimizer needs Python for the scripts in its folder and the command-line tools its instructions call (python3, pip and just). Our summary lists: Python 3.

Does RAG Accuracy Optimizer access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is RAG Accuracy Optimizer safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does RAG Accuracy Optimizer use?

RAG Accuracy Optimizer is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does RAG Accuracy Optimizer use?

About 5.5k tokens (SKILL.md is roughly 22k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 29k tokens, read only when the agent opens those files.

What are the alternatives to RAG Accuracy Optimizer?

Skills that share tags, products or a category with RAG Accuracy Optimizer: Aliyun Opensearch Search (cinience/alicloud-skills, 397 stars), Tanyuan Search (infometa/workbuddyskills, 344 stars), Snowflake Cortex AI (Mindrally/skills, 268 stars) and Postgres (timescale/pg-aiguide, 1.9k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains RAG Accuracy Optimizer?

LeoYeAI (a GitHub user) maintains it in LeoYeAI/openclaw-master-skills, which has 2,159 GitHub stars. The repository holds 972 skills in this directory. The repository was last updated on July 20, 2026.

Source: LeoYeAI/openclaw-master-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.