Agent skill

RAG Troubleshoot

by lyonzin in lyonzin/knowledge-rag

When the user reports a bug, error message, stack trace, unexpected behavior, or "why is this broken" question, search the corpus first for prior occurrences, known fixes, or related runbooks.

MITAuto-check passedAI & LLM Engineering

Install RAG Troubleshoot

skills CLI
$ npx skills add lyonzin/knowledge-rag --skill rag-troubleshoot -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install lyonzin/knowledge-rag rag-troubleshoot --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/lyonzin/knowledge-rag.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/workflow/rag-troubleshoot .claude/skills/rag-troubleshoot && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
rag-troubleshoot
GitHub stars
292
Token cost
~1.8k tokens
SKILL.md length
476 words
Files
1
Skills in repo
10
Repo updated
First seen
Licence
MIT

At a glance

When the user reports a bug, error message, stack trace, unexpected behavior, or "why is this broken" question, search the corpus first for prior occurrences, known fixes, or related runbooks.

  • Works in 3 steps: The error signature itself — exception… → The affected component / function —… → Prior incidents with similar symptoms —…
  • Unexpected behavior
  • SKILL.md covers When to use this skill, What this skill commits to, Steps and Examples, plus 2 more sections
  • Needs ERR_INVALID_TOKEN

What it does

RAG Troubleshoot is an agent skill from lyonzin/knowledge-rag. When the user reports a bug, error message, stack trace, unexpected behavior, or "why is this broken" question, search the corpus first for prior occurrences, known fixes, or related runbooks. Prevents re-solving problems the team already solved. Trigger on any error signature, exception name, stack trace snippet, or "why does X fail" query.

Its SKILL.md is about 1.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering Retrieval-augmented generation, Debugging and Runbooks and postmortems. It works with Model Context Protocol. The repository describes itself as: Local RAG MCP server for Claude Code — hybrid search (semantic + BM25), cross-encoder reranking, 13 MCP tools, 20 format parsers. Zero external servers, zero API keys. The licence is MIT.

When your agent uses it

  • Unexpected behavior
  • Why is this broken question
  • Search the corpus first for prior occurrences
  • Related runbooks

Example prompts

  • “why is this broken”
  • “why does X fail”
  • “/rag-troubleshoot”

Requirements

  • Python 3
  • A credential in ERR_INVALID_TOKEN

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. The error signature itself — exception name, first line of stack, unique error code
  2. The affected component / function — module name, service name, feature
  3. Prior incidents with similar symptoms — even if the error message differs

What it can do on your machine

Read from SKILL.md and the folder at commit df9cccb. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • ERR_INVALID_TOKEN

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

RAG Troubleshoot loads about 1.8k tokens when it runs. Until then it costs about 90 tokens; SKILL.md has 476 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~90
When it runs · the whole SKILL.md, loaded when a task matches
~1.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from lyonzin/knowledge-rag at commit df9cccb, republished under its MIT licence (© lyonzin). 476 words, ~1,760 tokens.

Download SKILL.mdSave it as .claude/skills/rag-troubleshoot/SKILL.md (or your agent's skills folder).
name
rag-troubleshoot
description
When the user reports a bug, error message, stack trace, unexpected behavior, or "why is this broken" question, search the corpus first for prior occurrences, known fixes, or related runbooks. Prevents re-solving problems the team already solved. Trigger on any error signature, exception name, stack trace snippet, or "why does X fail" query.
metadata.type
rag-workflow
metadata.kind
workflow
metadata.target
any-mcp-client

rag-troubleshoot — RAG-first debugging

When to use this skill

Trigger the moment the user reports:

  • A specific error message, stack trace, or exception
  • "It broke", "it fails", "not working", "returns null when it shouldn't"
  • Unexpected behavior in a component
  • Alert / incident triage
  • "Why does X do Y" where Y is wrong
  • A pasted log line
  • A failed CI/CD run

The core insight: many bugs are already solved somewhere in your corpus — runbook, postmortem, incident report, prior fix commit, ADR, chat thread indexed via add_from_url. Search first.


What this skill commits to

Before proposing a fix, the agent searches for:

  1. The error signature itself — exception name, first line of stack, unique error code
  2. The affected component / function — module name, service name, feature
  3. Prior incidents with similar symptoms — even if the error message differs

Only after those three come back empty does the agent apply general debugging techniques.


Steps

  1. Extract error signatures from the user's message:

    • Exception class name (ValueError, ConnectionError, TimeoutError, etc.)
    • Error code (ERR_INVALID_TOKEN, E42_INDEX_MISS, HTTP status)
    • First distinctive line of the stack trace
    • Unique keywords ("segfault at ...", "cannot connect to ...")
  2. First search — exact error signature:

    search_knowledge(query="<exception name> <error keyword>", max_results=5, min_score=0.15)

    Example: search_knowledge(query="ConnectionError refused chromadb")

  3. Second search — affected component + symptom:

    search_knowledge(query="<component name> <what happens>", max_results=5)

    Example: search_knowledge(query="reindex fails halfway large corpus")

  4. Third search — historical incidents / postmortems:

    search_knowledge(query="incident postmortem <component>", max_results=5)

    Or with a category filter if your corpus categorizes by document type:

    search_knowledge(query="<component> failure", category="incidents")
  5. Evaluate:

    ResultAction
    Prior fix existsApply it, cite the source, verify it still applies
    Related runbook existsFollow the runbook, cite it
    Related incident, different fixRead the incident, adapt
    Nothing foundDo general-purpose debugging, but flag that "this may be a new failure mode"
  6. After solving, if the fix is non-obvious, chain to rag-index-decisions to add a short note back to the corpus so the next occurrence is one search away.


Show full SKILL.md (177 more words)Show less

Examples

Example 1 — Prior fix

User: "Getting EmbeddingModelLoadError: sticky failure when starting knowledge-rag. What's happening?"

Agent:

search_knowledge(query="EmbeddingModelLoadError sticky failure") → 3 hits, top 0.91

Reply: "This is documented in [CHANGELOG.md#v381]. The `sticky _load_failed` flag was
added in v3.8.1 as a hotfix — when the ONNX model fails to load, subsequent calls
re-raise immediately instead of looping through HF Hub download attempts. The fix is
in [providers/embeddings/fastembed.py:145]. Root cause is usually one of:
  1. HF Hub unreachable at first startup (offline / DNS)
  2. Corrupt cached model in models_cache/
  3. Disk full during model download

Try: (a) check network to huggingface.co, (b) rm -rf models_cache/ and restart,
(c) df -h. If none apply, this may be a new failure mode."

Example 2 — Runbook

User: "Our search results are empty after nuclear rebuild. What now?"

Agent:

search_knowledge(query="empty results nuclear rebuild") → 4 hits
1st hit: docs/reindex-operations.md (0.87)

get_document(filepath="docs/reindex-operations.md")
→ Full runbook including "post-rebuild sanity check" section

Reply: "Per [docs/reindex-operations.md#post-rebuild-sanity-check], after a nuclear
rebuild you should:
1. Confirm `get_index_stats().chunks_count` > 0
2. Run one `search_knowledge('any indexed keyword')` — should return hits
3. If empty, check for stale FTS5 marker (v4.8.2+ specific) — delete
   `data/fts5_migration.state` and restart.
The most common cause is exactly the FTS5 stale marker case, added in v4.8.3
[CHANGELOG.md#v483]."

Example 3 — Nothing found, general debug

User: "My Rust binding for the MCP server segfaults on second call. Ideas?"

Agent:

search_knowledge(query="rust binding segfault MCP") → 0 hits
search_knowledge(query="rust FFI second call crash") → 0 hits
search_knowledge(query="rust bindings segfault") → 0 hits

Reply: "Nothing in the local corpus about Rust bindings (searched 'rust binding
segfault MCP' + 2 paraphrases, 0 hits). This appears to be a new failure mode.
General-purpose debugging directions:
  1. Rebuild with `RUST_BACKTRACE=full` to get the panic location
  2. Check FFI lifetime — is a Python object dropped while Rust still holds a ref?
  3. Try with `--single-instance` mode to rule out concurrency
  4. Reduce to a minimal repro: 1 call, 2 calls, does it always fail on N=2?

Once you find the root cause, worth indexing back — see rag-index-decisions."

Edge cases

  • Very generic error ("KeyError") — 1st search will have low precision. Add the component name early or skip the exact-error search and start with component + symptom.
  • Multi-line stack traces — extract the deepest custom frame (not stdlib) as the search seed.
  • Error is intermittent / rare — still search, but weight the "prior incidents" query more heavily.
  • User pasted only the symptom, not the error — ask for the error message + a stack trace before searching. Guessing wastes search calls.

  • rag-check-first — the parent skill (troubleshooting is a specialized variant).
  • rag-cite-sources — when you propose a fix, cite the source that documented it.
  • rag-index-decisions — after solving a novel bug, index the postmortem for next time.
  • rag-web-fallback — for truly novel errors, escalate to GitHub / StackOverflow after RAG comes back empty.

© lyonzin, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/workflow/rag-troubleshoot of lyonzin/knowledge-rag.

Open the folder on GitHubat commit df9cccb

Compare with similar skills

RAG Troubleshoot next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

RAG Troubleshoot compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
RAG Troubleshoot this skilllyonzin/knowledge-rag292—~1.8kAutomated safety check: PassMIT
AutoRAG Setup and RepairMarker-Inc-Korea/AutoRAG5.1k—~5.6kAutomated safety check: PassMIT
Sciverseopendatalab/Sciverse-Agent-Tools120—~3kAutomated safety check: PassCustom licence
Pgvector Semantic Searchtimescale/pg-aiguide1.9k—~3.8kAutomated safety check: PassApache-2.0
Postgres Hybrid Text Searchtimescale/pg-aiguide1.9k—~3.1kAutomated safety check: PassApache-2.0
SynalinksSynaLinks/synalinks-skills907—~4.8kAutomated safety check: PassApache-2.0

Similar skills

  • AutoRAG Setup and Repair

    Marker-Inc-Korea/AutoRAG

    Installs, configures, and repairs AutoRAG's search model, approved folders, indexes, and datasources, and registers its Lite MCP server.

    5.1k GitHub stars~5.6k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Sciverse

    opendatalab/Sciverse-Agent-Tools

    A skill your agent uses when the user needs academic paper retrieval — searching scientific literature by author/year/journal, finding paper chunks for RAG-style citations, or expanding original…

    120 GitHub stars~3k tokensUpdated 21 days ago
    AI & LLM EngineeringAuto-check passed
  • Pgvector Semantic Search

    timescale/pg-aiguide

    A skill your agent uses for setting up vector similarity search with pgvector for AI/ML embeddings, RAG applications, or semantic search.

    1.9k GitHub stars~3.8k tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check passed
  • Postgres Hybrid Text Search

    timescale/pg-aiguide

    A skill your agent uses to implement hybrid search combining BM25 keyword search with semantic vector search using Reciprocal Rank Fusion (RRF).

    1.9k GitHub stars~3.1k tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check passed
  • Synalinks

    SynaLinks/synalinks-skills

    A skill your agent uses for anything involving the Synalinks neuro-symbolic LM framework (Keras-inspired): DataModel/Field/Input, JSON operators (+ & | ^ ~), synalinks.ops…

    907 GitHub stars~4.8k tokensUpdated 12 days ago
    AI & LLM EngineeringAuto-check passed
  • AutoRAG Doctor

    Marker-Inc-Korea/AutoRAG

    Diagnoses and repairs a broken AutoRAG install so every configured datasource is both indexed and returns real search hits.

    5.1k GitHub stars~3.4k tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from lyonzin/knowledge-rag

All 10 skills in this repo
  • RAG Check First

    lyonzin/knowledge-rag

    Before answering any technical question, code request, architecture decision, or factual claim, call searchknowledge to check the local corpus.

    292 GitHub stars~1.4k tokensUpdated yesterday
    Auto-check passed
  • RAG Cite Sources

    lyonzin/knowledge-rag

    Every technical claim drawn from the local corpus must ship with a source citation formatted as path:line or path:section.

    292 GitHub stars~1.4k tokensUpdated yesterday
    Auto-check passed
  • RAG Code Review

    lyonzin/knowledge-rag

    When performing code review on a PR, diff, snippet, or "look at this change" request, first consult the corpus for related ADRs, coding standards, prior patterns, and similar files.

    292 GitHub stars~1.8k tokensUpdated yesterday
    Auto-check passed
  • RAG Deep Dive

    lyonzin/knowledge-rag

    Three-step multi-tool workflow — search the corpus, fetch the most relevant document in full, then find similar documents.

    292 GitHub stars~1.6k tokensUpdated yesterday
    Auto-check passed
  • RAG Evaluate Quality

    lyonzin/knowledge-rag

    Measure retrieval quality using evaluateretrieval (MRR@5 and Recall@5) and getindexstats.

    292 GitHub stars~1.4k tokensUpdated yesterday
    Auto-check passed
  • RAG Index Decisions

    lyonzin/knowledge-rag

    After making a non-obvious architectural decision, solving a novel bug, agreeing on a coding standard, or reaching a conclusion worth remembering, index it back into the knowledge base so the next…

    292 GitHub stars~2k tokensUpdated yesterday
    Auto-check passed

Questions about RAG Troubleshoot

What does RAG Troubleshoot do?

When the user reports a bug, error message, stack trace, unexpected behavior, or "why is this broken" question, search the corpus first for prior occurrences, known fixes, or related runbooks. RAG Troubleshoot is an agent skill from lyonzin/knowledge-rag. When the user reports a bug, error message, stack trace, unexpected behavior, or "why is this broken" question, search the corpus first for prior occurrences, known fixes, or related runbooks.

When should I use RAG Troubleshoot?

RAG Troubleshoot fits situations like: unexpected behavior; why is this broken question; search the corpus first for prior occurrences; related runbooks.

How do I install RAG Troubleshoot in Claude Code?

Run `npx skills add lyonzin/knowledge-rag --skill rag-troubleshoot -a claude-code`. Or copy the skill folder (skills/workflow/rag-troubleshoot in lyonzin/knowledge-rag) into .claude/skills/rag-troubleshoot in your project. Claude Code loads it when a task matches its description.

How do I install RAG Troubleshoot in Codex?

Run `npx skills add lyonzin/knowledge-rag --skill rag-troubleshoot -a codex`. Or copy the skill folder (skills/workflow/rag-troubleshoot in lyonzin/knowledge-rag) into .agents/skills/rag-troubleshoot in your project. Codex loads it when a task matches its description.

Can I use RAG Troubleshoot in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add lyonzin/knowledge-rag --skill rag-troubleshoot -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/rag-troubleshoot, .gemini/skills/rag-troubleshoot, .github/skills/rag-troubleshoot and .opencode/skills/rag-troubleshoot in your project.

What does RAG Troubleshoot need to run?

Going by SKILL.md and its folder, RAG Troubleshoot needs credentials named ERR_INVALID_TOKEN. Our summary lists: Python 3; A credential in ERR_INVALID_TOKEN.

Does RAG Troubleshoot access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is RAG Troubleshoot safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does RAG Troubleshoot use?

RAG Troubleshoot is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does RAG Troubleshoot use?

About 1.8k tokens (SKILL.md is roughly 7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to RAG Troubleshoot?

Skills that share tags, products or a category with RAG Troubleshoot: AutoRAG Setup and Repair (Marker-Inc-Korea/AutoRAG, 5.1k stars), Sciverse (opendatalab/Sciverse-Agent-Tools, 120 stars), Pgvector Semantic Search (timescale/pg-aiguide, 1.9k stars) and Postgres Hybrid Text Search (timescale/pg-aiguide, 1.9k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains RAG Troubleshoot?

lyonzin (a GitHub user) maintains it in lyonzin/knowledge-rag, which has 292 GitHub stars. The repository holds 10 skills in this directory. The repository was last updated on October 9, 2026.

Source: lyonzin/knowledge-rag on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.