Agent skill

RAG Web Fallback

by lyonzin in lyonzin/knowledge-rag

Only reach for external web search when the local corpus comes back empty or clearly insufficient.

MITAuto-check passedAI & LLM Engineering

Install RAG Web Fallback

skills CLI
$ npx skills add lyonzin/knowledge-rag --skill rag-web-fallback -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install lyonzin/knowledge-rag rag-web-fallback --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/lyonzin/knowledge-rag.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/workflow/rag-web-fallback .claude/skills/rag-web-fallback && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
rag-web-fallback
GitHub stars
290
Token cost
~1.4k tokens
SKILL.md length
514 words
Files
1
Skills in repo
10
Repo updated
First seen
Licence
MIT

At a glance

Only reach for external web search when the local corpus comes back empty or clearly insufficient.

  • Works in 3 steps: Attempt at least one search_knowledge… → Attempt a second call with alternative… → Only then, if genuinely no local…
  • Tasks that involve Retrieval-augmented generation
  • SKILL.md covers When to use this skill, What this skill commits to, Steps and Examples, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

RAG Web Fallback is an agent skill from lyonzin/knowledge-rag. Only reach for external web search when the local corpus comes back empty or clearly insufficient. Forces the agent to try knowledge-rag first, then explicitly document why it needed to escalate. Prevents wasted API cost, latency, and (in air-gapped deployments) accidental network calls.

Its SKILL.md is about 1.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering Retrieval-augmented generation and Web search. It works with Model Context Protocol. The repository describes itself as: Local RAG MCP server for Claude Code — hybrid search (semantic + BM25), cross-encoder reranking, 13 MCP tools, 20 format parsers. Zero external servers, zero API keys. The licence is MIT.

When your agent uses it

  • Tasks that involve Retrieval-augmented generation
  • Tasks that involve Web search

Example prompts

  • “/rag-web-fallback”

Requirements

  • Python 3

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Attempt at least one search_knowledge call before any web call, using appropriately extracted keywords.
  2. Attempt a second call with alternative phrasing if the first returned nothing relevant.
  3. Only then, if genuinely no local coverage, escalate to web search — with an explicit note to the user that RAG was checked and came back…

What it can do on your machine

Read from SKILL.md and the folder at commit df9cccb. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

RAG Web Fallback loads about 1.4k tokens when it runs. Until then it costs about 76 tokens; SKILL.md has 514 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~76
When it runs · the whole SKILL.md, loaded when a task matches
~1.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from lyonzin/knowledge-rag at commit df9cccb, republished under its MIT licence (© lyonzin). 514 words, ~1,407 tokens.

Download SKILL.mdSave it as .claude/skills/rag-web-fallback/SKILL.md (or your agent's skills folder).
name
rag-web-fallback
description
Only reach for external web search when the local corpus comes back empty or clearly insufficient. Forces the agent to try knowledge-rag first, then explicitly document why it needed to escalate. Prevents wasted API cost, latency, and (in air-gapped deployments) accidental network calls.
metadata.type
rag-workflow
metadata.kind
workflow
metadata.target
any-mcp-client

rag-web-fallback — local first, web only when necessary

When to use this skill

Whenever the agent is tempted to call an external tool for information (WebSearch, WebFetch, mcp__exa__*, mcp__context7__*, etc.), route through this skill's checklist first.

Applies to:

  • Any request that could be answered from indexed docs before hitting the network
  • Any factual claim the agent is about to look up externally
  • Any "give me the latest on X" question where X might be in the corpus

Does NOT apply to:

  • Deliberate freshness requests: "what happened today", "current news"
  • User explicitly asks for a web search
  • Live infra queries (external API calls that are the actual task, not for context)

What this skill commits to

The agent will:

  1. Attempt at least one search_knowledge call before any web call, using appropriately extracted keywords.
  2. Attempt a second call with alternative phrasing if the first returned nothing relevant.
  3. Only then, if genuinely no local coverage, escalate to web search — with an explicit note to the user that RAG was checked and came back empty.

Steps

  1. First-pass RAG:

    search_knowledge(query="<extracted keywords>", max_results=5, min_score=0.15)
  2. If 0 hits (or top score < 0.3), try a paraphrase. Common variants:

    • Swap synonyms ("auth" ↔ "authentication", "reindex" ↔ "rebuild", "ttl" ↔ "expiration")
    • Broaden ("payment refund cancellation" → "refund process")
    • Narrow ("architecture" → "authentication architecture")
  3. Second-pass RAG:

    search_knowledge(query="<paraphrased keywords>", max_results=5, min_score=0.10)

    Note: lower min_score (0.10) to catch weaker matches.

  4. Decide:

    Local result qualityAction
    Solid hit (score > 0.5)Answer from RAG, do NOT hit web
    Weak hit (score 0.15–0.5)Answer from RAG and add a web-search suggestion "for more context"
    No hitExplicitly tell the user "corpus does not have this" then optionally invoke web search
  5. If escalating to web, prefix the answer with:

    "The local knowledge base does not cover this (searched <query1>, <query2>, 0 relevant hits). Falling back to external web search."

    Then run the web tool (WebSearch, mcp__exa__web_search_exa, etc.).


Show full SKILL.md (210 more words)Show less

Examples

Example 1 — RAG covers it, no web needed

User: "What is our retry policy for external HTTP calls?"

Agent:

search_knowledge(query="retry policy external http") → 2 hits, top score 0.79

Reply: "Per [docs/adr/0031-retries.md], the policy is 3 attempts with exponential
backoff starting at 1s, max 30s. See services/common/retry.py:12–28 for the impl."

(No web search performed.)

Example 2 — RAG empty, web fallback explicit

User: "What are the top CVEs in Python cryptography package in Q3 2026?"

Agent:

search_knowledge(query="cve python cryptography 2026", min_score=0.10) → 0 hits
search_knowledge(query="python cryptography vulnerabilities", min_score=0.10) → 0 hits

Reply: "The local corpus does not track CVE data (searched 'cve python cryptography 2026'
and 'python cryptography vulnerabilities', 0 hits). Falling back to external web search..."

[Then calls WebSearch or mcp__cti__vuln_cve]

Example 3 — Weak local hit + web supplement

User: "Explain the LangChain LCEL syntax."

Agent:

search_knowledge(query="langchain lcel syntax") → 1 hit, score 0.22
(a passing mention in one of your ADRs comparing frameworks)

Reply: "The corpus has one passing mention of LangChain LCEL in
[docs/adr/0044-framework-choice.md]. That does not fully explain the syntax
though — for the syntax reference itself, I will pull from LangChain's public docs."

[Then calls WebFetch on the LangChain docs URL]

Edge cases

  • Air-gapped deployment — some enterprise setups have NO web tools enabled at all. In that case, this skill collapses to "search RAG; if empty, tell the user the corpus does not have it and stop." Do not fabricate.
  • User is impatient / one-shot query — you can shorten the 2-pass check to 1 pass. Do not skip it entirely.
  • Freshness-critical query ("what is the current CVE score for CVE-2024-1234") — skip RAG; the corpus is likely stale on live data. But state that you are skipping and why.
  • Related MCP tools available — if the workspace has mcp__cti__*, mcp__shodan__*, mcp__virustotal__*, mcp__context7__*, etc., prefer those over generic web search — they are usually more precise and faster.

  • rag-check-first — the prerequisite (this skill is check-first with an explicit escalation rule).
  • rag-cite-sources — if you DO answer from RAG, cite. If from web, cite the URL.
  • rag-troubleshoot — troubleshooting has its own escalation path (StackOverflow, GitHub issues) that follows the same pattern.

© lyonzin, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/workflow/rag-web-fallback of lyonzin/knowledge-rag.

Open the folder on GitHubat commit df9cccb

Compare with similar skills

RAG Web Fallback next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

RAG Web Fallback compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
RAG Web Fallback this skilllyonzin/knowledge-rag290—~1.4kAutomated safety check: PassMIT
Local RAG Searchnkapila6/mcp-local-rag1341 repos~1.6kAutomated safety check: PassMIT
Digoal Read Think Writerdigoal/blog8.6k—~1.3kAutomated safety check: PassGPL-2.0
MCP Local RAGshinpr/mcp-local-rag407—~4.4kAutomated safety check: PassMIT
Pgvector Semantic Searchtimescale/pg-aiguide1.9k1 repos~3.8kAutomated safety check: PassApache-2.0
AutoRAG Setup and RepairMarker-Inc-Korea/AutoRAG5.1k—~5.3kAutomated safety check: PassMIT

Similar skills

  • Local RAG Search

    nkapila6/mcp-local-rag

    Efficiently perform web searches using the mcp-local-rag server with semantic similarity ranking.

    134 GitHub starsUsed in 1 repo~1.6k tokens
    AI & LLM EngineeringAuto-check passed
  • 以 digoal/德哥 的第一人称口吻重写一篇文章。流程是:先吃透原文(必要时用 mcpMiniMaxwebsearch 拓展资料库),再用德哥的语气重新讲一遍,输出 markdown 到当前项目的 markdown/ 目录(SVG 图存到 markdown/svg/,文中以 ![描述](svg/xxx.svg)…

    8.6k GitHub stars~1.3k tokensUpdated 10 days ago
    AI & LLM EngineeringAuto-check passed
  • MCP Local RAG

    shinpr/mcp-local-rag

    Searches, saves, and maintains a local document index through a local RAG MCP server.

    407 GitHub stars~4.4k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Pgvector Semantic Search

    timescale/pg-aiguide

    A skill your agent uses for setting up vector similarity search with pgvector for AI/ML embeddings, RAG applications, or semantic search.

    1.9k GitHub starsUsed in 1 repo~3.8k tokens
    AI & LLM EngineeringAuto-check passed
  • AutoRAG Setup and Repair

    Marker-Inc-Korea/AutoRAG

    Installs, configures, and repairs AutoRAG's search model, approved folders, indexes, and datasources, and registers its Lite MCP server.

    5.1k GitHub stars~5.3k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Sciverse

    opendatalab/Sciverse-Agent-Tools

    A skill your agent uses when the user needs academic paper retrieval — searching scientific literature by author/year/journal, finding paper chunks for RAG-style citations, or expanding original…

    119 GitHub stars~3k tokensUpdated 18 days ago
    AI & LLM EngineeringAuto-check passed

More from lyonzin/knowledge-rag

All 10 skills in this repo
  • RAG Check First

    lyonzin/knowledge-rag

    Before answering any technical question, code request, architecture decision, or factual claim, call searchknowledge to check the local corpus.

    290 GitHub stars~1.4k tokensUpdated 4 days ago
    Auto-check passed
  • RAG Cite Sources

    lyonzin/knowledge-rag

    Every technical claim drawn from the local corpus must ship with a source citation formatted as path:line or path:section.

    290 GitHub stars~1.4k tokensUpdated 4 days ago
    Auto-check passed
  • RAG Code Review

    lyonzin/knowledge-rag

    When performing code review on a PR, diff, snippet, or "look at this change" request, first consult the corpus for related ADRs, coding standards, prior patterns, and similar files.

    290 GitHub stars~1.8k tokensUpdated 4 days ago
    Auto-check passed
  • RAG Deep Dive

    lyonzin/knowledge-rag

    Three-step multi-tool workflow — search the corpus, fetch the most relevant document in full, then find similar documents.

    290 GitHub stars~1.6k tokensUpdated 4 days ago
    Auto-check passed
  • RAG Troubleshoot

    lyonzin/knowledge-rag

    When the user reports a bug, error message, stack trace, unexpected behavior, or "why is this broken" question, search the corpus first for prior occurrences, known fixes, or related runbooks.

    290 GitHub stars~1.8k tokensUpdated 4 days ago
    Auto-check passed
  • RAG Evaluate Quality

    lyonzin/knowledge-rag

    Measure retrieval quality using evaluateretrieval (MRR@5 and Recall@5) and getindexstats.

    290 GitHub stars~1.4k tokensUpdated 4 days ago
    Auto-check passed

Questions about RAG Web Fallback

What does RAG Web Fallback do?

Only reach for external web search when the local corpus comes back empty or clearly insufficient. RAG Web Fallback is an agent skill from lyonzin/knowledge-rag. Only reach for external web search when the local corpus comes back empty or clearly insufficient.

When should I use RAG Web Fallback?

RAG Web Fallback fits situations like: tasks that involve Retrieval-augmented generation; tasks that involve Web search.

How do I install RAG Web Fallback in Claude Code?

Run `npx skills add lyonzin/knowledge-rag --skill rag-web-fallback -a claude-code`. Or copy the skill folder (skills/workflow/rag-web-fallback in lyonzin/knowledge-rag) into .claude/skills/rag-web-fallback in your project. Claude Code loads it when a task matches its description.

How do I install RAG Web Fallback in Codex?

Run `npx skills add lyonzin/knowledge-rag --skill rag-web-fallback -a codex`. Or copy the skill folder (skills/workflow/rag-web-fallback in lyonzin/knowledge-rag) into .agents/skills/rag-web-fallback in your project. Codex loads it when a task matches its description.

Can I use RAG Web Fallback in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add lyonzin/knowledge-rag --skill rag-web-fallback -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/rag-web-fallback, .gemini/skills/rag-web-fallback, .github/skills/rag-web-fallback and .opencode/skills/rag-web-fallback in your project.

What does RAG Web Fallback need to run?

SKILL.md names no scripts, command-line tools or credentials: RAG Web Fallback is instructions for the agent only. Our summary lists: Python 3.

Does RAG Web Fallback access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is RAG Web Fallback safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does RAG Web Fallback use?

RAG Web Fallback is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does RAG Web Fallback use?

About 1.4k tokens (SKILL.md is roughly 5.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to RAG Web Fallback?

Skills that share tags, products or a category with RAG Web Fallback: Local RAG Search (nkapila6/mcp-local-rag, 134 stars), Digoal Read Think Writer (digoal/blog, 8.6k stars), MCP Local RAG (shinpr/mcp-local-rag, 407 stars) and Pgvector Semantic Search (timescale/pg-aiguide, 1.9k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains RAG Web Fallback?

lyonzin (a GitHub user) maintains it in lyonzin/knowledge-rag, which has 290 GitHub stars. The repository holds 10 skills in this directory. The repository was last updated on October 4, 2026.

Source: lyonzin/knowledge-rag on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.