Agent skill

RAG Onboard Context

by lyonzin in lyonzin/knowledge-rag

At the start of every new session or when the topic shifts significantly, probe the knowledge base to learn what is indexed.

MITAuto-check passedAI & LLM Engineering

Install RAG Onboard Context

skills CLI
$ npx skills add lyonzin/knowledge-rag --skill rag-onboard-context -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install lyonzin/knowledge-rag rag-onboard-context --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/lyonzin/knowledge-rag.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/foundation/rag-onboard-context .claude/skills/rag-onboard-context && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
rag-onboard-context
GitHub stars
292
Token cost
~1.5k tokens
SKILL.md length
546 words
Files
1
Skills in repo
10
Repo updated
First seen
Licence
MIT

At a glance

At the start of every new session or when the topic shifts significantly, probe the knowledge base to learn what is indexed.

  • Works in 6 steps: Get index health → Enumerate categories → Probe 1–2 topics the user is likely to… → …
  • Tasks that involve Retrieval-augmented generation
  • SKILL.md covers When to use this skill, What this skill commits to, Steps and Examples, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

RAG Onboard Context is an agent skill from lyonzin/knowledge-rag. At the start of every new session or when the topic shifts significantly, probe the knowledge base to learn what is indexed. Calls getindexstats + listcategories + a couple of exploratory searchknowledge queries. Prevents the agent from operating blind or making wrong assumptions about what the corpus contains.

Its SKILL.md is about 1.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering Retrieval-augmented generation and Knowledge bases. It works with Model Context Protocol. The repository describes itself as: Local RAG MCP server for Claude Code — hybrid search (semantic + BM25), cross-encoder reranking, 13 MCP tools, 20 format parsers. Zero external servers, zero API keys. The licence is MIT.

When your agent uses it

  • Tasks that involve Retrieval-augmented generation
  • Tasks that involve Knowledge bases

Example prompts

  • “/rag-onboard-context”

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Get index health
  2. Enumerate categories
  3. Probe 1–2 topics the user is likely to touch. If the user's first message mentions a domain, probe it. Otherwise, probe the top 2 largest…
  4. Optionally, if you need concrete file names, call
  5. Store the summary internally — do not necessarily surface it to the user unless they ask. The value is that YOU now know
  6. From here on, rag-check-first handles every subsequent request with this context in mind.

What it can do on your machine

Read from SKILL.md and the folder at commit df9cccb. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

RAG Onboard Context loads about 1.5k tokens when it runs. Until then it costs about 84 tokens; SKILL.md has 546 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~84
When it runs · the whole SKILL.md, loaded when a task matches
~1.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from lyonzin/knowledge-rag at commit df9cccb, republished under its MIT licence (© lyonzin). 546 words, ~1,538 tokens.

Download SKILL.mdSave it as .claude/skills/rag-onboard-context/SKILL.md (or your agent's skills folder).
name
rag-onboard-context
description
At the start of every new session or when the topic shifts significantly, probe the knowledge base to learn what is indexed. Calls get_index_stats + list_categories + a couple of exploratory search_knowledge queries. Prevents the agent from operating blind or making wrong assumptions about what the corpus contains.
metadata.type
rag-workflow
metadata.kind
foundation
metadata.target
any-mcp-client

rag-onboard-context — know your corpus before you use it

When to use this skill

Run this skill:

  • At the start of any new conversation where knowledge-rag is available and the user is about to ask substantive questions
  • When the topic shifts significantly (from security to infrastructure, from dev to research)
  • After a major reindex (the corpus content changed under you)
  • When the user says "I just indexed new docs" or similar

Do NOT run repeatedly — once per session is usually enough. The query cache applies to searches; it does not cache every inspection tool.


What this skill commits to

Before diving into task-specific work, the agent gathers a mental map of the corpus:

  • How big — chunk count, document count, cache health
  • What is in it — which categories exist, roughly how many docs per category
  • What flavor — probe 1–2 topics the user is likely to ask about, see if there are hits

That mental map informs every subsequent decision: whether to search first, whether to fall back to the web, how to phrase queries.


Steps

  1. Get index health:

    get_index_stats()

    Read the nested stats object: total_documents, total_chunks, query_cache.hit_rate, embedding_model, and embedding_dim.

  2. Enumerate categories:

    list_categories()

    Note: which categories exist and how many docs each has. This tells you the corpus's shape.

  3. Probe 1–2 topics the user is likely to touch. If the user's first message mentions a domain, probe it. Otherwise, probe the top 2 largest categories:

    search_knowledge(query="<category-typical-term>", max_results=3, snippet_mode=true)
  4. Optionally, if you need concrete file names, call:

    list_documents(category="<top-category>")

    This tool has no limit argument and returns the whole category. For a large category, prefer a targeted search_knowledge call with max_results=3.

  5. Store the summary internally — do not necessarily surface it to the user unless they ask. The value is that YOU now know:

    • "The corpus has 15K chunks across 8 categories"
    • "Biggest categories are security (5K), development (3K), research (2K)"
    • "Search for mitre attack returned 3 solid hits — MITRE content is well indexed"
    • "Search for kubernetes returned nothing — try synonyms or inspect relevant categories before concluding the topic is absent"
  6. From here on, rag-check-first handles every subsequent request with this context in mind.


Show full SKILL.md (197 more words)Show less

Examples

Example 1 — Fresh session, dev context

User: "Hi. I want to add a new REST endpoint."

Agent (internal, before answering):

get_index_stats() → 12,043 chunks / 3,214 docs, cache hit 0%, model bge-small-en-v1.5
list_categories() → categories {development: 1200, security: 800, general: 1169, adrs: 45}
search_knowledge(query="REST endpoint", max_results=3) → 3 hits, top score 0.71

Now I know: this project has ADRs indexed. I should look at those before designing the endpoint.

Reply: "Let me check how the existing endpoints are structured..."
[calls search_knowledge with more specific queries]

Example 2 — Fresh session, security context

User: "We have a suspected LSASS dump alert. Can you help triage?"

Agent (internal):

get_index_stats() → 45,000 chunks / 5,800 docs
list_categories() → categories {redteam: 2400, blueteam: 1800, ctf: 1000, mitre: 600}
search_knowledge(query="LSASS credential dump", max_results=5) → 5 hits

Now I know: this is a security-focused RAG. LSASS is well-covered. Full triage is
going to be corpus-driven.

Reply: "Pulling triage steps from the runbooks and MITRE mapping..."

Example 3 — Empty / near-empty corpus

Agent (internal):

get_index_stats() → 12 chunks / 3 docs, cache hit 0%
list_categories() → {general: 3}

Signal: this is a fresh install. The corpus is nearly empty. Do NOT
"rag-check-first" aggressively — most queries will miss. Behave more
like a normal (RAG-optional) assistant and gently suggest the user
run `reindex_documents` after adding their docs.

Reply: "I notice your knowledge base only has 3 documents indexed. Before I dive in,
would you like to point me at your docs folder so we can index them first?"

Edge cases

  • Very large corpus (>100K chunks) — keep probes bounded and measure response time. First-use embedding/BM25 initialization and available RAM can affect latency.
  • Cache already used — stats.query_cache.hit_rate > 0 describes hits during this process lifetime. It does not prove an entry for the next query remains valid; the cache is in memory and has a TTL.
  • User immediately asks a task-specific question — do onboarding silently in the background and continue answering. Do not stall the user with a "let me look around first" message unless the corpus is empty.
  • Categories are empty ({}) — inspect index counts and indexing errors. An empty category map alone does not identify the cause. Search without a category filter while keeping max_results bounded.

  • rag-check-first — the workhorse skill that runs on every subsequent turn, informed by what onboarding revealed.
  • rag-deep-dive — chained after check-first when a topic needs more depth.
  • rag-evaluate-quality — periodic checkup (weekly, not per-session).

© lyonzin, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/foundation/rag-onboard-context of lyonzin/knowledge-rag.

Open the folder on GitHubat commit df9cccb

Compare with similar skills

RAG Onboard Context next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

RAG Onboard Context compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
RAG Onboard Context this skilllyonzin/knowledge-rag292—~1.5kAutomated safety check: PassMIT
AI Learning JournalLeoYeAI/openclaw-master-skills2.2k—~2.6kAutomated safety check: PassMIT
Amazon Bedrockaws/agent-toolkit-for-aws2.8k—~8.6kAutomated safety check: PassApache-2.0
Gnogmickel/gno1151 repos~1.6kAutomated safety check: PassMIT
Gnogmickel/gno115—~11kAutomated safety check: PassMIT
MCP Local RAGshinpr/mcp-local-rag412—~4.4kAutomated safety check: PassMIT

Similar skills

  • AI Learning Journal

    LeoYeAI/openclaw-master-skills

    AI 学习记录与成长追踪工具。用于记录 AI/LLM 学习笔记、使用心得、Prompt 技巧、工具体验等,并提供学习指导和规划。当用户提到以下任何话题时都应使用此 skill:AI 学习记录、学习笔记、AI 使用心得、Prompt 工程学习、模型对比体验、AI 工具使用记录、LLM 学习、RAG 学习、Agent 学习、MCP 学习、AI 微调实践、AI 学习规划、怎么学 AI、AI…

    2.2k GitHub stars~2.6k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed
  • Amazon Bedrock

    aws/agent-toolkit-for-aws

    Official

    Builds generative AI applications on Amazon Bedrock. An agent skill from aws/agent-toolkit-for-aws.

    2.8k GitHub stars~8.6k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Gno

    gmickel/gno

    Search local documents, files, notes, and knowledge bases. An agent skill from gmickel/gno.

    115 GitHub starsUsed in 1 repo~1.6k tokens
    Knowledge ManagementAuto-check passed
  • Gno

    gmickel/gno

    Search local documents, files, notes, and knowledge bases. An agent skill from gmickel/gno.

    115 GitHub stars~11k tokensUpdated 2 days ago
    Knowledge ManagementAuto-check passed
  • MCP Local RAG

    shinpr/mcp-local-rag

    Searches, saves, and maintains a local document index through a local RAG MCP server.

    412 GitHub stars~4.4k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Local RAG Search

    nkapila6/mcp-local-rag

    Efficiently perform web searches using the mcp-local-rag server with semantic similarity ranking.

    134 GitHub starsUsed in 1 repo~1.6k tokens
    AI & LLM EngineeringAuto-check passed

More from lyonzin/knowledge-rag

All 10 skills in this repo
  • RAG Check First

    lyonzin/knowledge-rag

    Before answering any technical question, code request, architecture decision, or factual claim, call searchknowledge to check the local corpus.

    292 GitHub stars~1.4k tokensUpdated yesterday
    Auto-check passed
  • RAG Cite Sources

    lyonzin/knowledge-rag

    Every technical claim drawn from the local corpus must ship with a source citation formatted as path:line or path:section.

    292 GitHub stars~1.4k tokensUpdated yesterday
    Auto-check passed
  • RAG Code Review

    lyonzin/knowledge-rag

    When performing code review on a PR, diff, snippet, or "look at this change" request, first consult the corpus for related ADRs, coding standards, prior patterns, and similar files.

    292 GitHub stars~1.8k tokensUpdated yesterday
    Auto-check passed
  • RAG Deep Dive

    lyonzin/knowledge-rag

    Three-step multi-tool workflow — search the corpus, fetch the most relevant document in full, then find similar documents.

    292 GitHub stars~1.6k tokensUpdated yesterday
    Auto-check passed
  • RAG Troubleshoot

    lyonzin/knowledge-rag

    When the user reports a bug, error message, stack trace, unexpected behavior, or "why is this broken" question, search the corpus first for prior occurrences, known fixes, or related runbooks.

    292 GitHub stars~1.8k tokensUpdated yesterday
    Auto-check passed
  • RAG Evaluate Quality

    lyonzin/knowledge-rag

    Measure retrieval quality using evaluateretrieval (MRR@5 and Recall@5) and getindexstats.

    292 GitHub stars~1.4k tokensUpdated yesterday
    Auto-check passed

Questions about RAG Onboard Context

What does RAG Onboard Context do?

At the start of every new session or when the topic shifts significantly, probe the knowledge base to learn what is indexed. RAG Onboard Context is an agent skill from lyonzin/knowledge-rag. At the start of every new session or when the topic shifts significantly, probe the knowledge base to learn what is indexed.

When should I use RAG Onboard Context?

RAG Onboard Context fits situations like: tasks that involve Retrieval-augmented generation; tasks that involve Knowledge bases.

How do I install RAG Onboard Context in Claude Code?

Run `npx skills add lyonzin/knowledge-rag --skill rag-onboard-context -a claude-code`. Or copy the skill folder (skills/foundation/rag-onboard-context in lyonzin/knowledge-rag) into .claude/skills/rag-onboard-context in your project. Claude Code loads it when a task matches its description.

How do I install RAG Onboard Context in Codex?

Run `npx skills add lyonzin/knowledge-rag --skill rag-onboard-context -a codex`. Or copy the skill folder (skills/foundation/rag-onboard-context in lyonzin/knowledge-rag) into .agents/skills/rag-onboard-context in your project. Codex loads it when a task matches its description.

Can I use RAG Onboard Context in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add lyonzin/knowledge-rag --skill rag-onboard-context -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/rag-onboard-context, .gemini/skills/rag-onboard-context, .github/skills/rag-onboard-context and .opencode/skills/rag-onboard-context in your project.

What does RAG Onboard Context need to run?

SKILL.md names no scripts, command-line tools or credentials: RAG Onboard Context is instructions for the agent only.

Does RAG Onboard Context access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is RAG Onboard Context safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does RAG Onboard Context use?

RAG Onboard Context is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does RAG Onboard Context use?

About 1.5k tokens (SKILL.md is roughly 6.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to RAG Onboard Context?

Skills that share tags, products or a category with RAG Onboard Context: AI Learning Journal (LeoYeAI/openclaw-master-skills, 2.2k stars), Amazon Bedrock (aws/agent-toolkit-for-aws, 2.8k stars), Gno (gmickel/gno, 115 stars) and Gno (gmickel/gno, 115 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains RAG Onboard Context?

lyonzin (a GitHub user) maintains it in lyonzin/knowledge-rag, which has 292 GitHub stars. The repository holds 10 skills in this directory. The repository was last updated on October 9, 2026.

Source: lyonzin/knowledge-rag on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.