Agent skill

Nodetool RAG Indexing

by nodetool-ai in nodetool-ai/nodetool

Build NodeTool document ingestion, vector indexing, retrieval, and RAG pipelines.

AGPL-3.0Auto-check passedAI & LLM Engineering

Install Nodetool RAG Indexing

skills CLI
$ npx skills add nodetool-ai/nodetool --skill nodetool-rag-indexing -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install nodetool-ai/nodetool nodetool-rag-indexing --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/nodetool-ai/nodetool.git skills-src && mkdir -p .claude/skills && cp -r skills-src/packages/system-skills/nodetool-rag-indexing .claude/skills/nodetool-rag-indexing && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
nodetool-rag-indexing
GitHub stars
560
Token cost
~1.4k tokens
SKILL.md length
440 words
Files
1
Skills in repo
127
Repo updated
First seen
Licence
AGPL-3.0

At a glance

Build NodeTool document ingestion, vector indexing, retrieval, and RAG pipelines.

  • Tasks that involve Retrieval-augmented generation
  • SKILL.md covers Chunk Size Guidance, Index and Query
  • Calls curl; needs CHROMA_TOKEN
  • Tasks that involve Vector databases

What it does

Nodetool RAG Indexing is an agent skill from nodetool-ai/nodetool. Build NodeTool document ingestion, vector indexing, retrieval, and RAG pipelines.

Its SKILL.md is about 1.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering Retrieval-augmented generation and Vector databases. The repository describes itself as: Agent-first Creative Workspace. The licence is AGPL-3.0.

When your agent uses it

  • Tasks that involve Retrieval-augmented generation
  • Tasks that involve Vector databases

Example prompts

  • “/nodetool-rag-indexing”

Requirements

  • A credential in CHROMA_TOKEN

What it can do on your machine

Read from SKILL.md and the folder at commit 339f069. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • curl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use curl, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • CHROMA_TOKEN

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Nodetool RAG Indexing loads about 1.4k tokens when it runs. Until then it costs about 26 tokens; SKILL.md has 440 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~26
When it runs · the whole SKILL.md, loaded when a task matches
~1.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from nodetool-ai/nodetool at commit 339f069, republished under its AGPL-3.0 licence (© nodetool-ai). 440 words, ~1,426 tokens.

Download SKILL.mdSave it as .claude/skills/nodetool-rag-indexing/SKILL.md (or your agent's skills folder).
name
nodetool-rag-indexing
description
Build NodeTool document ingestion, vector indexing, retrieval, and RAG pipelines.
featured
true

You help users build Retrieval-Augmented Generation (RAG) pipelines in NodeTool.

RAG Architecture

INDEXING:  Documents → Load → Split → Embed → Store (vector collection)
QUERY:     Question → Embed → Search → Format → LLM → Answer

Vector Store Backends

NodeTool's vector store (@nodetool-ai/vectorstore) is backend-pluggable. The workflow nodes are the same regardless of backend — you pick the backend via configuration.

BackendBest forNotes
SQLite-vecDefault, local, embeddedNo external service
ChromaDBSelf-host / remoteCHROMA_URL, CHROMA_PATH, CHROMA_TOKEN
PineconeManaged cloudAPI-key based
Supabase (pgvector)Postgres-backedPairs with Supabase auth/storage

There are no FAISS nodes; the backends above cover local and hosted use.

Vector Nodes (vector.*)

All RAG nodes live under the single vector.* namespace (not vector.chroma.* or vector.faiss.*).

NodePurpose
vector.CollectionReference/select a collection by name (the collection ref other nodes consume)
vector.IndexTextChunkIndex a single text chunk with its embedding
vector.IndexStringIndex a string value
vector.IndexAggregatedTextIndex aggregated text
vector.IndexEmbeddingIndex a precomputed embedding
vector.IndexImageIndex an image
vector.QueryTextVector similarity search over text
vector.QueryImageVector similarity search over images
vector.HybridSearchVector + keyword search (best accuracy)
vector.GetDocumentsRetrieve specific documents
vector.CountCount documents in a collection
vector.PeekPreview collection contents
vector.RemoveOverlapDe-duplicate overlapping chunks in results

Query nodes (QueryText, QueryImage, HybridSearch) output ids, documents, metadatas, and distances (HybridSearch also returns scores).

Document Loading & Splitting

NodeNamespacePurpose
Codenodetool.codeEnumerate files with await workspace.list(dir)
LoadDocumentFilenodetool.documentLoad a PDF/TXT/MD into a document
Chunknodetool.textFixed-size word chunking with overlap (general purpose)
RegexSplitnodetool.textStructure-aware splitting on a delimiter pattern

Chunk Size Guidance

Content typeChunk sizeOverlap
Technical docs200-500 tokens50 tokens
Prose/articles300-600 tokens75 tokens
Code100-300 tokens25 tokens
Show full SKILL.md (190 more words)Show less

Indexing — Workflow Pattern

ListFiles → LoadDocumentFile → Chunk → IndexTextChunk(collection)

Pair every index/query node with a vector.Collection node (or a collection name) so they target the same store. Use the same embedding model for indexing and querying.

Indexing — HTTP API

bash
# Index a file into a collection
curl -X POST http://localhost:7777/api/collections/<name>/index \
  -H "Authorization: Bearer TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"file_path": "/path/to/document.pdf"}'

The server resolves the collection, runs its ingestion workflow if one is registered, otherwise falls back to split → embed → store.

Query — Workflow Pattern

ChatInput → HybridSearch(collection, top_k) → FormatText → Agent → Output
NodePurpose
ChatInputUser question
vector.HybridSearchVector + keyword retrieval (best accuracy)
vector.QueryTextVector-only retrieval (faster)
FormatTextBuild the context string for the LLM
AgentGenerate the answer from context + question
OutputReturn the answer

Complete RAG Example

Index

ListFiles("/docs/") → LoadDocumentFile → Chunk(length=400, overlap=50)
                                                  ↓
                                  IndexTextChunk(collection="my-docs")

Query

ChatInput("What is...?") → HybridSearch(collection="my-docs", top_k=5)
                                  ↓
                         FormatText(template="Context:\n{documents}\n\nQuestion: {query}")
                                  ↓
                         Agent(model=gpt-5.4, system="Answer using only the context provided.")
                                  ↓
                         Output

Environment Variables

bash
# ChromaDB backend (only when using Chroma — SQLite-vec needs no config)
CHROMA_URL=                          # Remote Chroma URL (empty = local)
CHROMA_PATH=~/.local/share/nodetool/chroma  # Local storage path
CHROMA_TOKEN=                        # Optional auth token

The embedding model is chosen on the index/query nodes via model selection (e.g. text-embedding-3-small, or a local sentence-transformers model).

Common Pitfalls

  • Embedding model mismatch: use the same embedding model for indexing and search.
  • Chunks too large: dilute the LLM context — keep to 200-500 tokens.
  • Chunks too small: sentences get fragmented; use 10-20% overlap.
  • Empty collection: index before querying — an unindexed collection returns nothing.
  • Wrong node names: it's vector.IndexTextChunk / vector.QueryText / vector.HybridSearch, not IndexTextChunks / TextSearch, and there is no vector.chroma.*/vector.faiss.* namespace.
  • No nodetool collections CLI: manage collections through the editor UI or the /api/collections/... endpoints.

© nodetool-ai, AGPL-3.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in packages/system-skills/nodetool-rag-indexing of nodetool-ai/nodetool.

Open the folder on GitHubat commit 339f069

Compare with similar skills

Nodetool RAG Indexing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Nodetool RAG Indexing compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Nodetool RAG Indexing this skillnodetool-ai/nodetool560—~1.4kAutomated safety check: PassAGPL-3.0
Chroma Vector DatabaseOrchestra-Research/AI-Research-SKILLs13k7 repos~2.3kAutomated safety check: PassMIT
Pgvector Semantic Searchtimescale/pg-aiguide1.9k—~3.8kAutomated safety check: PassApache-2.0
Postgres Hybrid Text Searchtimescale/pg-aiguide1.9k—~3.1kAutomated safety check: PassApache-2.0
RAG Implementationwshobson/agents40k9 repos~1.1kAutomated safety check: PassMIT
Vector DBRightNow-AI/openfang18k—~1kAutomated safety check: PassApache-2.0

Similar skills

  • Chroma Vector Database

    Orchestra-Research/AI-Research-SKILLs

    Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.

    13k GitHub starsUsed in 7 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Pgvector Semantic Search

    timescale/pg-aiguide

    A skill your agent uses for setting up vector similarity search with pgvector for AI/ML embeddings, RAG applications, or semantic search.

    1.9k GitHub stars~3.8k tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check passed
  • Postgres Hybrid Text Search

    timescale/pg-aiguide

    A skill your agent uses to implement hybrid search combining BM25 keyword search with semantic vector search using Reciprocal Rank Fusion (RRF).

    1.9k GitHub stars~3.1k tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check passed
  • RAG Implementation

    wshobson/agents

    Build retrieval-augmented generation systems: pick a vector database and embedding model, choose retrieval and reranking strategies, and start from a LangGraph pipeline.

    40k GitHub starsUsed in 9 repos~1.1k tokens
    AI & LLM EngineeringAuto-check passed
  • Vector DB

    RightNow-AI/openfang

    Vector database expert for embeddings, similarity search, RAG patterns, and indexing strategies

    18k GitHub stars~1k tokensUpdated 3 mo ago
    AI & LLM EngineeringAuto-check passed
  • Hunt RAG Vector

    elementalsouls/Claude-BugHunter

    Hunt vector-store / embedding-layer weaknesses in RAG pipelines (OWASP LLM08 Vector and Embedding Weaknesses) — persistent corpus poisoning that survives across sessions and users (distinct from…

    4.8k GitHub stars~2.6k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed

More from nodetool-ai/nodetool

All 127 skills in this repo
  • Beat Sync Editing

    nodetool-ai/nodetool

    Cut a NodeTool timeline to music and shape its pacing — detect the beat grid, place cuts on phrases, pick a cut type, build speed ramps with time remap, and give the piece an arc.

    560 GitHub stars~2.6k tokensUpdated today
    Auto-check passed
  • Caption Titles

    nodetool-ai/nodetool

    Add and animate a consistent text layer on an existing NodeTool timeline.

    560 GitHub stars~1.9k tokensUpdated today
    Auto-check passed
  • Color Motion

    nodetool-ai/nodetool

    Choose and animate colour on a NodeTool timeline, including shape and text gradients, colour grades, 3D LUTs, and dither.

    560 GitHub stars~2.3k tokensUpdated today
    Auto-check passed
  • Commercial Beat Sheet

    nodetool-ai/nodetool

    Write a shootable, precisely timed commercial beat sheet and store it as a NodeTool storyboard, with a consistent entity roster behind every shot.

    560 GitHub stars~4.6k tokensUpdated today
    Auto-check passed
  • Elevenlabs Audio Prompting

    nodetool-ai/nodetool

    Direct ElevenLabs speech, dialogue, sound effects and music — the bracketed audio tags v3 acts on and why the voice decides whether a tag lands, stability as the delivery dial, punctuation instead…

    560 GitHub stars~1.9k tokensUpdated today
    Auto-check passed
  • Frame Composition

    nodetool-ai/nodetool

    Stage the frame on a NodeTool timeline — grids, focal placement, safe areas per aspect ratio, depth layers and parallax, camera moves, and where elements enter and leave.

    560 GitHub stars~3.5k tokensUpdated today
    Auto-check passed

Questions about Nodetool RAG Indexing

What does Nodetool RAG Indexing do?

Build NodeTool document ingestion, vector indexing, retrieval, and RAG pipelines. Nodetool RAG Indexing is an agent skill from nodetool-ai/nodetool. Build NodeTool document ingestion, vector indexing, retrieval, and RAG pipelines.

When should I use Nodetool RAG Indexing?

Nodetool RAG Indexing fits situations like: tasks that involve Retrieval-augmented generation; tasks that involve Vector databases.

How do I install Nodetool RAG Indexing in Claude Code?

Run `npx skills add nodetool-ai/nodetool --skill nodetool-rag-indexing -a claude-code`. Or copy the skill folder (packages/system-skills/nodetool-rag-indexing in nodetool-ai/nodetool) into .claude/skills/nodetool-rag-indexing in your project. Claude Code loads it when a task matches its description.

How do I install Nodetool RAG Indexing in Codex?

Run `npx skills add nodetool-ai/nodetool --skill nodetool-rag-indexing -a codex`. Or copy the skill folder (packages/system-skills/nodetool-rag-indexing in nodetool-ai/nodetool) into .agents/skills/nodetool-rag-indexing in your project. Codex loads it when a task matches its description.

Can I use Nodetool RAG Indexing in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add nodetool-ai/nodetool --skill nodetool-rag-indexing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/nodetool-rag-indexing, .gemini/skills/nodetool-rag-indexing, .github/skills/nodetool-rag-indexing and .opencode/skills/nodetool-rag-indexing in your project.

What does Nodetool RAG Indexing need to run?

Going by SKILL.md and its folder, Nodetool RAG Indexing needs the command-line tools its instructions call (curl) and credentials named CHROMA_TOKEN. Our summary lists: A credential in CHROMA_TOKEN.

Does Nodetool RAG Indexing access the network?

SKILL.md contains no URLs. Its commands use curl, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Nodetool RAG Indexing safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Nodetool RAG Indexing use?

Nodetool RAG Indexing is published under the AGPL-3.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Nodetool RAG Indexing use?

About 1.4k tokens (SKILL.md is roughly 5.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Nodetool RAG Indexing?

Skills that share tags, products or a category with Nodetool RAG Indexing: Chroma Vector Database (Orchestra-Research/AI-Research-SKILLs, 13k stars), Pgvector Semantic Search (timescale/pg-aiguide, 1.9k stars), Postgres Hybrid Text Search (timescale/pg-aiguide, 1.9k stars) and RAG Implementation (wshobson/agents, 40k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Nodetool RAG Indexing?

nodetool-ai (a GitHub organization) maintains it in nodetool-ai/nodetool, which has 560 GitHub stars. The repository holds 127 skills in this directory. The repository was last updated on October 10, 2026.

Source: nodetool-ai/nodetool on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.