Ingest user-selected documents and retrieve cited procedures, standards, and design evidence from the local RAG knowledge base.

Apache-2.0Auto-check passedAI & LLM Engineering

Install RAG

skills CLI
$ npx skills add automateyournetwork/netclaw --skill rag -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install automateyournetwork/netclaw rag --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/automateyournetwork/netclaw.git skills-src && mkdir -p .claude/skills && cp -r skills-src/workspace/skills/rag .claude/skills/rag && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
rag
GitHub stars
676
Token cost
~2.8k tokens
SKILL.md length
1,369 words
Files
1
Skills in repo
120
Repo updated
First seen
Licence
Apache-2.0

At a glance

Ingest user-selected documents and retrieve cited procedures, standards, and design evidence from the local RAG knowledge base.

  • Works in 5 steps: Route first (Adaptive RAG) → Rewrite and decompose → Critique after every retrieval (Self-RAG) → …
  • Document questions
  • SKILL.md covers Overview, MCP Tools, The Four Knowledge Sources… and Workflow: Slack Document Upload, plus 6 more sections
  • Calls python3

What it does

RAG is an agent skill from automateyournetwork/netclaw. Ingest user-selected documents and retrieve cited procedures, standards, and design evidence from the local RAG knowledge base. Use for document questions, corpus management, or explicitly requested snapshots.

Its SKILL.md is about 2.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering Retrieval-augmented generation and Knowledge bases. The repository describes itself as: An AI agent that claws through your network. The licence is Apache-2.0.

When your agent uses it

  • Document questions
  • Corpus management
  • Explicitly requested snapshots

Example prompts

  • “/rag”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Route first (Adaptive RAG)
  2. Rewrite and decompose
  3. Critique after every retrieval (Self-RAG)
  4. Respect the iteration budget — 3 rounds per sub-query
  5. Honest miss

What it can do on your machine

Read from SKILL.md and the folder at commit 95bb17e. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

RAG loads about 2.8k tokens when it runs. Until then it costs about 53 tokens; SKILL.md has 1,369 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~53
When it runs · the whole SKILL.md, loaded when a task matches
~2.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from automateyournetwork/netclaw at commit 95bb17e, republished under its Apache-2.0 licence (© automateyournetwork). 1,369 words, ~2,771 tokens.

Download SKILL.mdSave it as .claude/skills/rag/SKILL.md (or your agent's skills folder).
name
rag
description
Ingest user-selected documents and retrieve cited procedures, standards, and design evidence from the local RAG knowledge base. Use for document questions, corpus management, or explicitly requested snapshots.

Skill: RAG Knowledge Base

Purpose: Give NetClaw a fully offline, user-curated document knowledge base — vendor guides, standards (RFC/IEEE/vendor), customer design documents, install guides — with agentic retrieval, mandatory citations, and opt-in point-in-time snapshots.

Overview

Users teach NetClaw by uploading documents (Slack attachment, HUD Knowledge panel, or URL). Documents are parsed, chunked structure-aware, embedded locally, and stored at ~/.openclaw/rag/. Retrieval is a tool NetClaw invokes on its own judgment — iteratively, with self-critique — never a fixed pipeline.

This is NOT memory. The knowledge base holds only what users deliberately put into it. NetClaw's own experience (facts, session summaries, decisions, entity graphs) lives in the Memory MCP (memory_* tools, ~/.openclaw/memory/). Neither store writes into the other.

MCP Tools

ToolWHEN to use
rag_ingestA document file on disk should be learned
rag_ingest_base64A Slack attachment should be learned (decode → ingest)
rag_ingest_urlThe user asks to ingest a web page (ALWAYS preview scope first)
rag_searchQuestion concerns vendor procedures, customer standards, install steps, or ingested content
rag_listUser asks what the knowledge base contains
rag_statsUser asks about corpus size/health or retrieval telemetry
rag_update_metadataFix a document's doc_type/title/version
rag_deleteUser asks to remove a document (CONFIRM with the user first)
rag_reindexChunking/embedding config changed (CONFIRM with the user first)
rag_snapshotUser EXPLICITLY asks to store live output for later comparison (confirm scope first — never automatic)

The Four Knowledge Sources (routing rules)

Route every question to the right source. Most questions need NO retrieval.

  1. Parametric knowledge — timeless networking fundamentals (OSPF LSA types, BGP path selection). Answer directly. Do not search.
  2. Memory MCP (memory_recall, memory_get_facts, memory_get_decisions) — NetClaw's own past sessions, learned facts, and decisions about THIS network. "What was that BGP issue last month?" goes here, never to rag_search.
  3. RAG knowledge base (rag_search) — user-uploaded documents. Vendor procedures, customer standards, install steps. Check it BEFORE declaring ignorance on these topics.
  4. Live MCP servers (pyATS, NetBox, etc.) — current network state. NEVER answer a live-state question from the RAG store. The only exception is an explicitly requested snapshot, whose age must always be shown.

When a question legitimately touches both Memory and the knowledge base, consult both — but attribute every part of the answer to its actual source. Memory recall is never presented as a document citation, and vice versa.

"Remember this document" → ingestion (RAG). "Remember that PE2 is in maintenance until Friday" → Memory (memory_record_fact).

Workflow: Slack Document Upload

When a user posts a file attachment in the Slack channel and asks NetClaw to learn it:

  1. Download the attachment content.

  2. Call the ingest tool with the base64-encoded file:

    bash
    python3 $MCP_CALL "python3 -u $RAG_MCP_SCRIPT" rag_ingest_base64 \
      '{"filename": "wlc-9800-upgrade-guide.pdf", "content_base64": "<b64>", "doc_type": "vendor"}'
  3. Infer doc_type from the user's words ("this is our customer standard" → customer); default other.

  4. Confirm in-thread with what was learned: title, doc_type, page count, chunk count, collection — plus the example_question from the response, so the user learns what they can now ask.

  5. If the response is deduplicated: true, say the document was already indexed. If reindexed: true, say the prior version was replaced.

  6. Errors (unsupported format, size cap, parse failure) are reported verbatim — never silently swallowed.

Supported formats: PDF, Markdown, HTML, TXT, DOCX, XLSX, PPTX, VSDX natively; legacy DOC/XLS/PPT/VSD when LibreOffice is installed.

Workflow: URL Ingestion (always preview first)

When a user asks to ingest a web page:

  1. Preview: rag_ingest_url {"url": "...", "mode": "preview"} — returns the page title, the same-domain pages it links to (depth 1, capped at RAG_CRAWL_MAX_PAGES), and a scope_token. No ingestion happens.
  2. Confirm scope with the user: "That page links to 6 same-domain pages. Ingest just the page, or all 7?" ALWAYS offer the single-page fallback. If the preview was truncated, say so.
  3. Ingest:
    • Single page: rag_ingest_url {"url": "...", "mode": "ingest"}
    • Page + linked pages (only after explicit confirmation): rag_ingest_url {"url": "...", "mode": "ingest", "include_linked": true, "scope_token": "<from preview>"}
  4. Confirm as with file ingestion. Each page records its own URL as source. PDFs served at URLs are handled automatically.

Never call mode="ingest" with include_linked=true without having shown the preview and obtained the user's confirmation — the server rejects a missing/stale scope_token.

The Agentic Retrieval Protocol

Retrieval is a tool you wield, not a pipeline you sit inside. For every question:

1. Route first (Adaptive RAG)

Decide whether to retrieve AT ALL using the four-source rules above. Most questions need no retrieval:

QuestionRouteWhy
"What's the OSPF LSA type for external routes?"Answer directlyTimeless fundamental
"What was that BGP issue last month on PE2?"memory_recallYour own past experience
"What does our customer standard require for change windows?"rag_search (filter doc_type: customer)User-uploaded document
"What's the current BGP state on PE2?"Live MCP (pyATS)Live network state — NEVER RAG
"Upgrade the lab WLC per our standards"rag_search for procedure/standards, THEN live MCP for pre-checksMixed — each part to its source
2. Rewrite and decompose

Rewrite conversational phrasing into retrieval-friendly queries ("how do I get the new code on the WLC" → "WLC software upgrade procedure install activate commit"). Decompose multi-part questions into independent sub-queries, retrieved separately with scoped filters, then synthesized. Give each sub-query a sub_query_id and pass round on every rag_search call so the budget is auditable.

Show full SKILL.md (531 more words)Show less
3. Critique after every retrieval (Self-RAG)

Grade the returned chunks: do they actually answer the question?

  • Yes → answer, citing every claim.
  • Partially / no → rewrite the query (different terms, add filters) and re-retrieve.
  • low_confidence: true on results → treat as signal, never as answer material.
4. Respect the iteration budget — 3 rounds per sub-query

Each sub-query gets at most 3 retrieval rounds (initial + 2 refinements; RAG_MAX_ROUNDS). Stop condition: when a sub-query's budget is exhausted, stop retrieving for it and state plainly what you found (cited) and what remains unanswered. Do not loop. Do not pad.

5. Honest miss

If the corpus doesn't cover the topic — empty results, corpus_empty: true, or only low-confidence chunks after refinement — say so:

"The knowledge base doesn't cover Nexus 9300 upgrades. I can ingest a document if you have one — drop it in this channel."

NEVER answer from irrelevant or low-confidence chunks. NEVER present a guess as knowledge-base content. A fabricated answer is worse than an honest gap.

Citation Rules (non-negotiable)

Every claim derived from retrieved content carries a citation:

[WLC 9800 Upgrade Guide §4.2, p.31 — ingested 2026-07-01]

Use the citation field returned by rag_search verbatim. A claim you cannot attribute to a specific retrieved chunk is NOT presented as coming from the knowledge base. Chunk IDs stay in the retrieval log — never show them to users.

When synthesizing across multiple documents, every combined claim must be supportable by at least one cited chunk. Do not blend unrelated chunks into a claim none of them supports — that is synthesis hallucination.

Snapshots — the ONLY sanctioned RAG use of live data

rag_snapshot exists for one purpose: explicit point-in-time comparison ("snapshot the BGP tables so we can compare next month"). It is NEVER a substitute for a live MCP query.

Absolute prohibition: NEVER invoke rag_snapshot automatically, from a heartbeat, on a schedule, or as a side effect of another task. It runs only on an explicit human request, and only after you confirm scope.

Workflow:

  1. User explicitly asks to snapshot live output.
  2. Confirm scope before executing (HIIL gate): which devices, which commands, what label. "Scope: PE1, PE2 — show ip bgp via pyATS, into snapshot_core-bgp_<timestamp>. Confirm?"
  3. Collect the output via the existing MCP tools (pyATS etc.).
  4. Call rag_snapshot {"label": "core-bgp", "content": "<output>", "source_description": "core router BGP tables", "devices": ["PE1","PE2"], "commands": ["show ip bgp"]}.
  5. Report what was stored including the per-type redaction counts from the response — secrets (passwords, SNMP communities, auth keys, pre-shared keys) are scrubbed before vectorization, and a zero count is reported explicitly.

When ANY answer later uses snapshot data:

  • Lead with the staleness notice: "From snapshot core-bgp — captured 2026-07-16 14:02 UTC — 31 days ago (live state is available via MCP tools): …" Use the age_human/staleness_notice fields verbatim.
  • Remind the user that live state is one MCP call away.
  • Snapshots older than RAG_SNAPSHOT_WARN_DAYS (default 90) are flagged stale — surface the flag. They are never auto-deleted; offer deletion instead.

Environment Variables

VariableDefaultPurpose
RAG_DATA_DIR~/.openclaw/ragPersistent store (never ~/.openclaw/memory/)
RAG_EMBEDDING_MODELBAAI/bge-small-en-v1.5Local embedding model
RAG_RERANKER_MODELcross-encoder/ms-marco-MiniLM-L-6-v2Local reranker
RAG_RERANK_ENABLEDtrueDisable on low-resource hosts
RAG_MAX_DOC_MB / RAG_MAX_DOC_PAGES100 / 1000Per-document ingestion caps
RAG_CRAWL_MAX_PAGES30Depth-1 crawl preview bound
RAG_SNAPSHOT_WARN_DAYS90Snapshot staleness warning
RAG_MAX_ROUNDS3Retrieval rounds per sub-query
RAG_MCP_SCRIPT—Path to rag-mcp/rag_mcp_server.py for $MCP_CALL

Example Usage

text
User: (attaches customer-wlan-standard.docx) learn this, it's our customer standard
NetClaw: Learned "Customer WLAN Standard" (customer, 24 pages, 61 chunks, documents).
         Try asking: "What does our standard 'Customer WLAN Standard' say about maintenance windows?"

© automateyournetwork, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in workspace/skills/rag of automateyournetwork/netclaw.

Open the folder on GitHubat commit 95bb17e

Compare with similar skills

RAG next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

RAG compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
RAG this skillautomateyournetwork/netclaw676—~2.8kAutomated safety check: PassApache-2.0
Blockify Integrationiternal-technologies-partners/blockify-agentic-data-optimization315—~6.2kAutomated safety check: NotesCustom licence
Agentsop Difyagentsope/SkillAlchemy466—~5.4kAutomated safety check: NotesMIT
Penguin SDKPrism-Shadow/penguin-harness2.5k—~11kAutomated safety check: PassApache-2.0
Sc QAopen-edge-platform/edge-ai-suites140—~2.3kAutomated safety check: PassApache-2.0
RAG AssistantAtmosphere/atmosphere3.8k—~504Automated safety check: PassApache-2.0

Similar skills

  • Blockify Integration

    iternal-technologies-partners/blockify-agentic-data-optimization

    Process documents with Blockify API to create optimized IdeaBlocks for RAG.

    315 GitHub stars~6.2k tokensUpdated 5 mo ago
    AI & LLM EngineeringAuto-check: notes
  • Agentsop Dify

    agentsope/SkillAlchemy

    SOP for building LLM applications on Dify — visual workflow + chatflow + agent + RAG knowledge base + plugin marketplace + observability, self-hostable.

    466 GitHub stars~5.4k tokensUpdated today
    AI & LLM EngineeringAuto-check: notes
  • Penguin SDK

    Prism-Shadow/penguin-harness

    A skill your agent uses whenever the user wants to build an agent application — their own program with an embedded agent, such as an AI app, an agentic app or a RAG app.

    2.5k GitHub stars~11k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Sc QA

    open-edge-platform/edge-ai-suites

    Ask a natural-language question against indexed content via the Content Search RAG Q&A endpoint.

    140 GitHub stars~2.3k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • RAG Assistant

    Atmosphere/atmosphere

    Knowledge base assistant that retrieves and cites documents from a curated index.

    3.8k GitHub stars~504 tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Langchain4j RAG Implementation Patterns

    giuseppe-trisciuoglio/developer-kit

    Provides Retrieval-Augmented Generation (RAG) implementation patterns with LangChain4j for Java.

    356 GitHub stars~3.3k tokensUpdated 29 days ago
    AI & LLM EngineeringAuto-check: notes

More from automateyournetwork/netclaw

All 120 skills in this repo
  • EVE-NG Lab Topology Design

    automateyournetwork/netclaw

    Entry point for designing EVE-NG network labs: classifies the request, gathers missing requirements, proposes options and validates the resulting topology.

    676 GitHub stars~612 tokensUpdated 3 days ago
    Auto-check passed
  • ACI Policy Change Deployment

    automateyournetwork/netclaw

    Deploys Cisco ACI policy changes only behind an approved ServiceNow Change Request, capturing pre and post-change fault baselines and rolling back automatically on a fault delta.

    676 GitHub stars~4.2k tokensUpdated 3 days ago
    Auto-check passed
  • Cisco ACI Fabric Health Audit

    automateyournetwork/netclaw

    Runs a phased health audit of a Cisco ACI fabric through MCP tools: node status, links, tenant and policy review, faults and endpoint learning.

    676 GitHub stars~2.9k tokensUpdated 3 days ago
    Auto-check passed
  • Anta Validation

    automateyournetwork/netclaw

    Validate Arista EOS network state against ANTA's pre-built 208-test catalogue, with structured pass/fail verdicts.

    676 GitHub stars~1.2k tokensUpdated 3 days ago
    Auto-check passed
  • Arista Cvp

    automateyournetwork/netclaw

    Arista CloudVision Portal (CVP) automation via REST API — device inventory, events, connectivity monitoring, tag management (4 tools).

    676 GitHub stars~2.2k tokensUpdated 3 days ago
    Auto-check: notes
  • AWS Cloud Monitoring

    automateyournetwork/netclaw

    AWS CloudWatch monitoring — metrics, alarms, log queries, VPC flow log analysis, network performance.

    676 GitHub stars~1k tokensUpdated 3 days ago
    Auto-check passed

Questions about RAG

What does RAG do?

Ingest user-selected documents and retrieve cited procedures, standards, and design evidence from the local RAG knowledge base. RAG is an agent skill from automateyournetwork/netclaw. Ingest user-selected documents and retrieve cited procedures, standards, and design evidence from the local RAG knowledge base.

When should I use RAG?

RAG fits situations like: document questions; corpus management; explicitly requested snapshots.

How do I install RAG in Claude Code?

Run `npx skills add automateyournetwork/netclaw --skill rag -a claude-code`. Or copy the skill folder (workspace/skills/rag in automateyournetwork/netclaw) into .claude/skills/rag in your project. Claude Code loads it when a task matches its description.

How do I install RAG in Codex?

Run `npx skills add automateyournetwork/netclaw --skill rag -a codex`. Or copy the skill folder (workspace/skills/rag in automateyournetwork/netclaw) into .agents/skills/rag in your project. Codex loads it when a task matches its description.

Can I use RAG in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add automateyournetwork/netclaw --skill rag -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/rag, .gemini/skills/rag, .github/skills/rag and .opencode/skills/rag in your project.

What does RAG need to run?

Going by SKILL.md and its folder, RAG needs the command-line tools its instructions call (python3). Our summary lists: Python 3.

Does RAG access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is RAG safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does RAG use?

RAG is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does RAG use?

About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to RAG?

Skills that share tags, products or a category with RAG: Blockify Integration (iternal-technologies-partners/blockify-agentic-data-optimization, 315 stars), Agentsop Dify (agentsope/SkillAlchemy, 466 stars), Penguin SDK (Prism-Shadow/penguin-harness, 2.5k stars) and Sc QA (open-edge-platform/edge-ai-suites, 140 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains RAG?

automateyournetwork (a GitHub user) maintains it in automateyournetwork/netclaw, which has 676 GitHub stars. The repository holds 120 skills in this directory. The repository was last updated on October 5, 2026.

Source: automateyournetwork/netclaw on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.