Agent skill

Basemind Documents

by Goldziher in Goldziher/basemind

Semantic + full-text search over documents and the web via basemind's RAG store — PDFs, Office, HTML, email, images (OCR), plus scraped/crawled web pages, with cross-encoder reranking, keyword and…

MITAuto-check passedAI & LLM Engineering

Install Basemind Documents

skills CLI
$ npx skills add Goldziher/basemind --skill basemind-documents -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Goldziher/basemind basemind-documents --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Goldziher/basemind.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/basemind-documents .claude/skills/basemind-documents && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
basemind-documents
GitHub stars
106
Token cost
~1.2k tokens
SKILL.md length
482 words
Files
1
Skills in repo
10
Repo updated
First seen
Licence
MIT

At a glance

Semantic + full-text search over documents and the web via basemind's RAG store — PDFs, Office, HTML, email, images (OCR), plus scraped/crawled web pages, with cross-encoder reranking, keyword and…

  • Tasks that involve Retrieval-augmented generation
  • SKILL.md covers Requirements, Tool routing, What a hit carries and Examples, plus 1 more section
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Tasks that involve Search implementation

What it does

Basemind Documents is an agent skill from Goldziher/basemind. Semantic + full-text search over documents and the web via basemind's RAG store — PDFs, Office, HTML, email, images (OCR), plus scraped/crawled web pages, with cross-encoder reranking, keyword and named-entity (NER) filters, and per-document summaries. Reach for it whenever the user asks to "search the docs / PDFs", "find where a topic is discussed", "pull this URL into context", or "what does the documentation say about X".

Its SKILL.md is about 1.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering Retrieval-augmented generation, Search implementation and PDF. It works with Model Context Protocol. The repository describes itself as: Full AI context and content layer for coding agents over one MCP server — tree-sitter code-map, document RAG, shared memory, multi-agent comms, web crawl, git history + blame… The licence is MIT.

When your agent uses it

  • Tasks that involve Retrieval-augmented generation
  • Tasks that involve Search implementation
  • Tasks that involve PDF

Example prompts

  • “search the docs / PDFs”
  • “find where a topic is discussed”
  • “pull this URL into context”
  • “/basemind-documents”

What it can do on your machine

Read from SKILL.md and the folder at commit 2f6da31. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Basemind Documents loads about 1.2k tokens when it runs. Until then it costs about 112 tokens; SKILL.md has 482 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~112
When it runs · the whole SKILL.md, loaded when a task matches
~1.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Goldziher/basemind at commit 2f6da31, republished under its MIT licence (© Goldziher). 482 words, ~1,210 tokens.

Download SKILL.mdSave it as .claude/skills/basemind-documents/SKILL.md (or your agent's skills folder).
name
basemind-documents
description
Semantic + full-text search over documents and the web via basemind's RAG store — PDFs, Office, HTML, email, images (OCR), plus scraped/crawled web pages, with cross-encoder reranking, keyword and named-entity (NER) filters, and per-document summaries. Reach for it whenever the user asks to "search the docs / PDFs", "find where a topic is discussed", "pull this URL into context", or "what does the documentation say about X".

basemind-documents — document RAG and web ingestion

basemind extracts 90+ file formats (PDF, Office, HTML, email, images via OCR) into a LanceDB vector store and answers meaning-based queries with cross-encoder reranking. Web pages scraped or crawled into the same store are searchable the same way. This is the surface for "find the passage about X", not "grep for the string X".

basemind first, open-the-file fallback. Prefer memory mode documents over opening PDFs/Office/HTML by hand, and the web tools over ad-hoc fetching. For source code use basemind-code-search instead — this skill is for prose and documents.

Requirements

  • memory mode documents needs a build with --features documents (or full); the other memory modes need --features memory. Without them the tools dispatch but return an MCP error.
  • Web ingestion (web modes scrape / crawl / map) needs --features crawl. When that feature is off these tools are not registered at all — they simply won't appear in the tool list.
  • Documents must be scanned first: basemind scan with the documents feature extracts and embeds them into the machine-global cache (Linux ~/.local/share/basemind/, macOS ~/Library/Application Support/basemind/; override BASEMIND_DATA_HOME). See the basemind-scan skill.

Tool routing

QuestionMCP toolCLI
"Semantic search over PDFs/Office/HTML docs?"memory { mode: "documents", query: "…" }basemind memory documents "query"
"Narrow to docs mentioning an entity?"memory { mode: "documents", query: "…", entity_category: "…" }(MCP only)
"Narrow to docs with a keyword?"memory { mode: "documents", query: "…", keywords_contains: "…" }(MCP only)
"Filter by file type?"memory { mode: "documents", query: "…", mime_type: "application/pdf" }basemind memory documents "…" --mime-type application/pdf
"Pull a single URL into RAG?"web { mode: "scrape", url: "…" } (robots-aware)basemind web scrape <url>
"Ingest a docs site section?"web { mode: "crawl", url: "…" }basemind web crawl <seed-url>
"What URLs exist on this site?"web { mode: "map", url: "…" }basemind web map <url>
"Recall something the agent stored earlier?"memory mode get, list, or searchbasemind memory get "key" / list / search "q"
"Remember this for future sessions?"memory { mode: "put", key, value }basemind memory put "key" "value"
Show full SKILL.md (164 more words)Show less

What a hit carries

memory mode documents returns chunk-level hits with path, chunk_idx, the matched text, byte span, vector distance, and — when enabled at scan time — a cross-encoder rerank_score in [0,1], the parent document's keywords and named entities (NER), and a document-level summary. Use entity_category / keywords_contains to constrain to documents whose parent carries a matching entity or keyword (AND-combined when both are set).

Examples

text
memory { mode: "documents", query: "how is the index schema versioned", limit: 5 }
→ docs/architecture.pdf#chunk3  rerank 0.91  "INDEX_SCHEMA_VER reads from RELEASE_MINOR…"
  README.md#chunk12             rerank 0.74  "…wipe-on-mismatch rebuilds from source…"

web { mode: "crawl", url: "https://docs.example.com/guide" }
→ ingested 24 pages under scope "web:docs.example.com"

memory { mode: "documents", query: "rate limiting", mime_type: "text/html" }
→ web:docs.example.com/limits#chunk1  rerank 0.88  "requests are capped at …"

Notes

  • Crawled/scraped pages land in the documents table tagged with a scope of web:<host> (override in web mode scrape); memory mode documents searches every ingested document.
  • robots.txt is honoured by default; only [crawl].respect_robots_txt = false in the repo-root basemind.toml (config-file-only) disables it. The crawler SSRF-blocks private/loopback hosts unless [crawl].allow_private_network = true.
  • Memory is scoped by the normalised git origin URL, so clones of the same repo share stored entries and unrelated repos do not.
  • Lists are capped (limit, default 100, max 1000); use next_cursor → cursor to page.

For code structure see basemind-code-search; for git history see basemind-git-history; for agent coordination see basemind-comms.

© Goldziher, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/basemind-documents of Goldziher/basemind.

Open the folder on GitHubat commit 2f6da31

Compare with similar skills

Basemind Documents next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Basemind Documents compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Basemind Documents this skillGoldziher/basemind106—~1.2kAutomated safety check: PassMIT
Postgrestimescale/pg-aiguide1.9k—~941Automated safety check: PassApache-2.0
Scholar RAGjoshzyj/open-scholar-skill168—~7.4kAutomated safety check: NotesCustom licence
I3brycewang-stanford/Auto-Empirical-Research-Skills4.5k—~1.8kAutomated safety check: PassCustom licence
MCP Local RAGshinpr/mcp-local-rag407—~4.4kAutomated safety check: PassMIT
Local RAG Searchnkapila6/mcp-local-rag1341 repos~1.6kAutomated safety check: PassMIT

Similar skills

  • Postgres

    timescale/pg-aiguide

    A skill your agent uses for any PostgreSQL database work — table design, indexing, data types, constraints, extensions (pgvector, PostGIS, TimescaleDB), search, and migrations.

    1.9k GitHub stars~941 tokensUpdated yesterday
    DatabasesAuto-check passed
  • Scholar RAG

    joshzyj/open-scholar-skill

    Build and query a local vector database + GraphRAG over your entire reference library (Zotero or a PDF folder) for literature review.

    168 GitHub stars~7.4k tokensUpdated 20 days ago
    Research & ScienceAuto-check: notes
  • I3

    brycewang-stanford/Auto-Empirical-Research-Skills

    RAG Builder with Parallel Document Processing Vector database construction with local embeddings (zero cost) Handles PDF download, text extraction, chunking, and vector database creation Absorbed B5…

    4.5k GitHub stars~1.8k tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check passed
  • MCP Local RAG

    shinpr/mcp-local-rag

    Searches, saves, and maintains a local document index through a local RAG MCP server.

    407 GitHub stars~4.4k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Local RAG Search

    nkapila6/mcp-local-rag

    Efficiently perform web searches using the mcp-local-rag server with semantic similarity ranking.

    134 GitHub starsUsed in 1 repo~1.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Pgvector Semantic Search

    timescale/pg-aiguide

    A skill your agent uses for setting up vector similarity search with pgvector for AI/ML embeddings, RAG applications, or semantic search.

    1.9k GitHub starsUsed in 1 repo~3.8k tokens
    AI & LLM EngineeringAuto-check passed

More from Goldziher/basemind

All 10 skills in this repo
  • Basemind Usage Stats

    Goldziher/basemind

    Shows a markdown dashboard of basemind activity in the session: tool call counts, top operations and estimated tokens saved against a grep and Read baseline.

    106 GitHub stars~558 tokensUpdated yesterday
    Auto-check passed
  • Basemind Code Context

    Goldziher/basemind

    Answers structural questions about a repository, such as where a symbol is defined, what calls it and what changed recently, through the basemind MCP server.

    106 GitHub stars~2.6k tokensUpdated yesterday
    Auto-check passed
  • Basemind Index Scan

    Goldziher/basemind

    Builds or refreshes the basemind code index from the command line, for when basemind reports no index or its MCP server is not running.

    106 GitHub stars~997 tokensUpdated yesterday
    Auto-check: warnings
  • Drives the basemind code-intelligence tool from the command line for symbol search, callers, git history and blame, document search and web crawling, sharing its cache with the MCP server.

    106 GitHub stars~2k tokensUpdated yesterday
    Auto-check passed
  • Basemind Code Search

    Goldziher/basemind

    Answers where code is defined, who calls it and what shape a file has from a pre-built index, returning paths and line numbers instead of file contents.

    106 GitHub stars~1.3k tokensUpdated yesterday
    Auto-check passed
  • Basemind Doctor

    Goldziher/basemind

    Diagnose and recover basemind when it isn't working — MCP tools missing or erroring, "no index" / "no indexed files", empty results that shouldn't be empty, or the MCP server seems dead.

    106 GitHub stars~1.1k tokensUpdated yesterday
    Auto-check passed

Questions about Basemind Documents

What does Basemind Documents do?

Semantic + full-text search over documents and the web via basemind's RAG store — PDFs, Office, HTML, email, images (OCR), plus scraped/crawled web pages, with cross-encoder reranking, keyword and…. Basemind Documents is an agent skill from Goldziher/basemind. Semantic + full-text search over documents and the web via basemind's RAG store — PDFs, Office, HTML, email, images (OCR), plus scraped/crawled web pages, with cross-encoder reranking, keyword and named-entity (NER) filters, and per-document summaries.

When should I use Basemind Documents?

Basemind Documents fits situations like: tasks that involve Retrieval-augmented generation; tasks that involve Search implementation; tasks that involve PDF.

How do I install Basemind Documents in Claude Code?

Run `npx skills add Goldziher/basemind --skill basemind-documents -a claude-code`. Or copy the skill folder (skills/basemind-documents in Goldziher/basemind) into .claude/skills/basemind-documents in your project. Claude Code loads it when a task matches its description.

How do I install Basemind Documents in Codex?

Run `npx skills add Goldziher/basemind --skill basemind-documents -a codex`. Or copy the skill folder (skills/basemind-documents in Goldziher/basemind) into .agents/skills/basemind-documents in your project. Codex loads it when a task matches its description.

Can I use Basemind Documents in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Goldziher/basemind --skill basemind-documents -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/basemind-documents, .gemini/skills/basemind-documents, .github/skills/basemind-documents and .opencode/skills/basemind-documents in your project.

What does Basemind Documents need to run?

SKILL.md names no scripts, command-line tools or credentials: Basemind Documents is instructions for the agent only.

Does Basemind Documents access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Basemind Documents safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Basemind Documents use?

Basemind Documents is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Basemind Documents use?

About 1.2k tokens (SKILL.md is roughly 4.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Basemind Documents?

Skills that share tags, products or a category with Basemind Documents: Postgres (timescale/pg-aiguide, 1.9k stars), Scholar RAG (joshzyj/open-scholar-skill, 168 stars), I3 (brycewang-stanford/Auto-Empirical-Research-Skills, 4.5k stars) and MCP Local RAG (shinpr/mcp-local-rag, 407 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Basemind Documents?

Goldziher (a GitHub user) maintains it in Goldziher/basemind, which has 106 GitHub stars. The repository holds 10 skills in this directory. The repository was last updated on October 8, 2026.

Source: Goldziher/basemind on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.