Agent skill

AutoRAG Setup and Repair

by Marker-Inc-Korea in Marker-Inc-Korea/AutoRAG

Installs, configures, and repairs AutoRAG's search model, approved folders, indexes, and datasources, and registers its Lite MCP server.

MITAuto-check passedAI & LLM Engineering

Install AutoRAG Setup and Repair

skills CLI
$ npx skills add Marker-Inc-Korea/AutoRAG --skill autorag-setup -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Marker-Inc-Korea/AutoRAG autorag-setup --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Marker-Inc-Korea/AutoRAG.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/autorag-setup .claude/skills/autorag-setup && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
autorag-setup
GitHub stars
5.1k
Token cost
~5.6k tokens
SKILL.md length
2,642 words
Files
1
Skills in repo
4
Repo updated
First seen
Licence
MIT

At a glance

Installs, configures, and repairs AutoRAG's search model, approved folders, indexes, and datasources, and registers its Lite MCP server.

  • Works in 5 steps: Probe every datasource for setup… → Auto-configure every datasource that… → Skip every datasource that probes… → …
  • Installing AutoRAG for the first time
  • SKILL.md covers Safety, Install the CLI if needed, Inspect existing configuration and Configure one search model, plus 8 more sections
  • Calls bun and npm; reaches openrouter.ai; needs OPENROUTER_API_KEY and GITHUB_TOKEN

What it does

This skill installs the CLI with Bun or npm if missing, then inspects existing configuration from a config file or environment variable, covering search paths, the chosen model, BM25 and sync settings, retrieval limits, and datasources, and preserves a working config and the user's explicit choices unless they ask to replace it or health checks fail.

It only ever inspects non-secret provider or model metadata and credential availability, storing environment-variable names rather than any actual key, and never scans the whole filesystem, indexes system or build directories, or moves, renames, or deletes source documents. For a model-free external-agent search setup it defers to a separate Lite-specific setup skill instead of configuring a search model itself.

When your agent uses it

  • Installing AutoRAG for the first time
  • Fixing a failed init, refresh, or health check
  • Adding a new folder or datasource to AutoRAG's index

Example prompts

  • “Install and configure AutoRAG for this project.”
  • “My AutoRAG indexes are stale, fix the health check failure.”
  • “Add this new folder as a datasource for AutoRAG.”

Requirements

  • Node.js 24 or newer, or Bun

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Probe every datasource for setup feasibility before asking the user
  2. Auto-configure every datasource that probes feasible — write its trusted
  3. Skip every datasource that probes infeasible (for example Notion or
  4. Set up a skipped datasource only when the user explicitly asks for it
  5. E-mail datasources (mail-export, mailcrawl) matter to most

What it can do on your machine

Read from SKILL.md and the folder at commit 29ccab9. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • bun
    • npm

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • openrouter.ai

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • OPENROUTER_API_KEY
    • GITHUB_TOKEN
    • TYPESAFE_API_KEY
    • AI_GATEWAY_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

AutoRAG Setup and Repair loads about 5.6k tokens when it runs. Until then it costs about 87 tokens; SKILL.md has 2,642 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~87
When it runs · the whole SKILL.md, loaded when a task matches
~5.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Marker-Inc-Korea/AutoRAG at commit 29ccab9, republished under its MIT licence (© Marker-Inc-Korea). 2,642 words, ~5,559 tokens.

Download SKILL.mdSave it as .claude/skills/autorag-setup/SKILL.md (or your agent's skills folder).
name
autorag-setup
description
Install and configure AutoRAG, register its Lite MCP server for external agents, or repair its single search model, approved roots, indexes, datasources, and health checks without exposing credentials. Use when AutoRAG or MCP is missing, init/refresh/health fails, indexes are stale, or the user wants to add folders or datasources.
license
MIT

AutoRAG setup

Use this skill when AutoRAG is unconfigured, the autorag CLI is missing, model resolution fails, indexes are missing or stale, or the user wants to change the document collection or datasources.

For model-free external-agent search, use autorag-lite-setup to configure and register autorag-mcp; do not configure a search model or install a search skill for Lite. This skill's model setup and CLI search apply to the full librarian, not the Lite MCP server.

Safety

  • Inspect only non-secret provider/model metadata and credential availability.
  • Never print, copy, migrate, compare, or persist credential values. Store only environment-variable names such as apiKeyEnv.
  • Do not scan the whole filesystem or home directory without explicit approval.
  • Never move, rename, edit, or delete source documents.
  • Do not index system trees, app bundles, caches, credential stores, node_modules, .git, dist, build, target, .cache, .autorag, or .jikji.

Install the CLI if needed

The CLI is @autorag/librarian (autorag). Runtime is Node.js ≥ 24 or Bun.

bash
command -v autorag >/dev/null || bun install -g @autorag/librarian
autorag --help

If Bun is unavailable, npm install -g @autorag/librarian is acceptable.

Inspect existing configuration

Check --config, AUTORAG_CONFIG, $AUTORAG_HOME/config.json, or ~/.autorag/config.json. Relevant fields are:

  • searchPaths, workspacePath, and memoryPath
  • model.provider, model.id, model.api, model.baseUrl, model.apiKeyEnv
  • bm25, minSync, and jikji
  • limits (retrieval, baseline-prefetch, and model-facing candidate caps)
  • datasources and ui

Preserve explicit user choices and a working config unless the user asks to replace them or health checks fail.

Configure one search model

AutoRAG is the specialized librarian agent. One configured model plans the search, calls retrieval and filesystem tools, reads sources, judges evidence, and curates the answer in one loop. There are no orchestrator/explorer roles.

Prefer a model with reliable tool calling and structured output, enough context for source excerpts, high output TPS, and low first-token latency. Use a larger reasoning model only when difficult synthesis or domain judgment matters more than latency.

Start from a provider/model the current runtime can actually call. A Pi-usable ChatGPT, Claude, Gemini, or other authenticated subscription is valid; an installed CLI or subscription that Pi cannot invoke is not. For custom OpenAI-compatible endpoints, record the real wire API, base URL, and credential environment-variable name.

Allowed api values:

  • openai-completions
  • openai-responses
  • anthropic-messages
  • openai-codex-responses
  • azure-openai-responses

If no callable setup can be established, ask the user for the provider, model id, API protocol, base URL when custom, and credential environment-variable name. Do not invent provider identities or model ids.

Propose and approve document roots

Explicit user paths win. Otherwise inspect only these likely document-dense candidates for existence and approximate supported-file counts, then ask for approval before indexing:

OSRecommendedOptional
macOS~/Documents, ~/Downloads, ~/Desktop~/Notes, ~/Obsidian, user-named project docs
LinuxXDG Documents/Downloads/Desktop or their ~/ defaults~/Notes, ~/Sync, Nextcloud/Syncthing roots
WindowsDocuments, Downloads, Desktop shell foldersOneDrive document roots, user-named project docs

Supported parsed formats are md, markdown, txt, text, pdf, docx, pptx, xlsx, xls, hwp, hwpx, and eml. OCR for jpg, jpeg, png, bmp, and tiff is optional (parserOptions.ocr.enabled). Do not present legacy .doc as a supported parsed format.

Keep the first-run set small, usually one to three roots. Present a concrete proposal and require yes, a narrowed keep-list, a custom list, or skip before running refresh.

Initialize

bash
autorag init \
  --search-paths "/path/to/documents,/path/to/notes" \
  --workspace "/path/to/workspace" \
  --model-provider PROVIDER \
  --model-id MODEL

Model resolution runs through the pi runtime. A configured provider/id is resolved against the pi runtime catalog — the built-in providers plus any ~/.pi/agent/models.json, custom, or extension providers — and credentials come from stored pi auth (~/.pi/agent/auth.json, API key or OAuth established with pi /login), then provider environment variables. When no model is configured, AutoRAG uses pi's defaultProvider/defaultModel when that provider has configured credentials; otherwise it falls back to the authenticated local codex runtime. So model flags may be omitted when either pi or the local runtime already supplies the intended model. Inspect what pi can resolve with autorag models list (--available shows only providers with configured credentials); it never prints credential values.

For a custom endpoint, add api, baseUrl, and apiKeyEnv to the single model object in the trusted config:

json
{
  "searchPaths": ["/path/to/documents"],
  "model": {
    "provider": "openrouter",
    "id": "anthropic/claude-sonnet-5.5",
    "api": "openai-completions",
    "baseUrl": "https://openrouter.ai/api/v1",
    "apiKeyEnv": "OPENROUTER_API_KEY"
  }
}

When provider/id names a pi runtime catalog model, the catalog entry stays the base: baseUrl, api, and any declared reasoning, input, contextWindow, or maxTokens override only those fields, and the catalog's reasoning, thinking, and compat settings are kept. Only an id outside the catalog (private proxy, Ollama, LiteLLM) gets a generic text model with a 128k context window unless those fields are declared.

Use --force only when intentionally replacing an existing config, and target it with an explicit path (--config or AUTORAG_CONFIG); --force refuses to replace the implicit ~/.autorag/config.json. Legacy cwd autorag.config.json is a migration source only and is never deleted by init.

Retrieval defaults

MinSync and Jikji are enabled by default. Leave them enabled unless the user explicitly asks otherwise. Indexing never happens while answering: a question only reads indexes that autorag refresh (or autorag watch) built, so run a refresh after setup and whenever documents change. A never-refreshed workspace still answers, but without MinSync/Jikji evidence.

Refresh (never a query) auto-installs the binaries: MinSync installs a verified GitHub release into <workspace>/.autorag/bin (minSync.autoInstall defaults to true), and Jikji installs jikji-cli through cargo (jikji.autoInstall defaults to true; requires the Rust toolchain). Set "autoInstall": false only when managing the binary yourself. Refresh is incremental: MinSync syncs only changed parsed mirrors, and jikji prepare reuses unchanged documents. Roots prepare in parallel.

Jikji stores its prepared corpus metadata in a hidden .jikji directory inside each indexed source root — that is Jikji's native index layout and intended behavior, not a misplaced artifact. Tell the user before the first refresh that <root>/.jikji will be created inside every approved document root (one per root, alongside the documents), and never delete or edit its contents; removing it only forces a full Jikji re-prepare on the next refresh.

Exact duplicate exclusion is enabled by default. AutoRAG invokes the external dupey CLI before parsed-mirror indexing, keeps the newest filesystem copy for each exact canonical-text hash, and excludes older copies from the mirror. Install dupey during setup when it is missing (command -v dupey || cargo install dupey --locked) and tell the user the feature is available; when installation is impossible, refresh continues without this optimization and the user is told duplicate exclusion is off. Set "excludeExactDuplicates": false to index every copy.

MinSync's default embedder is in-process native Qwen3 embeddings (native:Qwen/Qwen3-Embedding-0.6B, 1024 dimensions, MinSync 0.4.6). The default needs no embedder flags, no API key, and no external daemon (such as Ollama) — all embeddings run in-process locally and privately. For workspaces using the loopback llama-server gateway, prefetch the verified model with autorag models prefetch --profile qwen3-embedding-0.6b (or verify with autorag models verify --profile qwen3-embedding-0.6b).

The legacy Ollama/TEI adapter path (EmbeddingGemma, 768 dimensions) is supported only for existing legacy workspaces and manual QA; do not use it as a fresh install default.

Override the embedder only when intentionally using a different, for example remote, provider:

bash
autorag init \
  --embedder-id "voyageai/voyage-4-lite" \
  --embedder-base-url "https://openrouter.ai/api/v1" \
  --embedder-api-key-env "OPENROUTER_API_KEY" \
  --embedder-dimension 1024 \
  --embedder-batch-size 64

Only store the environment-variable name, never its value. Dimension and batch size must be positive integers, and the dimension must match the embedder (default Qwen3 is 1024; legacy EmbeddingGemma is 768; voyageai/voyage-4-lite via OpenRouter is 1024).

Jev routing and question decomposition (on by default)

Leave both enabled. They are the recommended setup: they make simple questions fast and multi-part questions thorough.

  • Jev (jev, default { "backend": "openrouter" }, model typesafe/jev-1.13) runs before the fast answer. It routes each question to local search, web search, a direct answer (general knowledge or small talk skips retrieval entirely), or the config branch, and decides whether to decompose it. On local search it also decides, per registered datasource, whether to search it before the fast answer, using each datasource's description and where similar past questions were answered (retrieval memory). When setting up a datasource, always write a description from what it actually holds: channels or rooms, people, topics, time range (for example "Team Slack, 2024-2026: #release and #on-call channels, dependabot notifications"). Jev is told descriptions are short, non-exhaustive summaries, so list the main content and do not try to list everything. After the fast answer it decides whether verification is needed, so a complete, evidence-backed fast answer ends the run.
  • Question decomposition (queryDecomposition, default model openrouter/qwen/qwen3.7-flash) splits a multi-part question into at most five search queries that run in parallel.

Both use the user's OPENROUTER_API_KEY; confirm it is set (test -n "$OPENROUTER_API_KEY", never print it) and tell the user Jev routing is on. Without the key, routing falls back to a single local search and the run always verifies, so searches still work. A query-route-fallback diagnostic (autorag search --debug) shows that state.

json
{
  "jev": { "backend": "openrouter" },
  "queryDecomposition": { "model": { "provider": "openrouter", "id": "qwen/qwen3.7-flash" } }
}

autorag init writes these defaults into new configs. To change them:

  • Jev backend: "backend": "typesafe" (TYPESAFE_API_KEY) or "vercel" (AI_GATEWAY_API_KEY).
  • Decomposition model: any catalog provider/id, with the same fields as the top-level model.
  • "queryDecomposition": false decomposes with the search model itself.
  • "jev": false turns routing off entirely. Do this only when the user explicitly opts out, for example because questions must never leave the machine (Jev and decomposition send the question text to OpenRouter).

When a user asks the running agent itself to change its settings (switch the default model, add a provider, check that a provider works), Jev's config branch loads this whole skill into that turn. The agent edits only the active config file (and models.json for a custom provider), verifies with autorag health --json and autorag models list --available, and reports each change as old → new through emit_autorag_results. It never prints a credential value and never uses init --force.

Show full SKILL.md (1,105 more words)Show less
Retrieval and ingest caps

limits bounds retrieval, baseline prefetch, and the candidate lists handed to the model. Every field is optional — an omitted field keeps the shipped default, so add the section only to tighten or widen a specific cap. Values must be positive integers; unknown keys (and unknown prefetch keys) fail config resolution.

FieldDefaultControls
mergedEvidenceCeiling500search_all_documents / model-free merge ceiling when the model omits topK
singleDatasourceTopK50search_datasource_* merge default when the model omits topK
minSyncTopK50MinSync semantic default topK
minSyncScopedQueryTopK100MinSync fetch cap applied when a scope narrows the query
toolDescriptionInstanceScopes8Instance scopes listed in one datasource tool description
prefetch.jikjiTopK30Jikji find candidate count
prefetch.minSyncTopK100MinSync retrieve candidate count
prefetch.jikjiPathLimit100Max Jikji answer paths rendered into the baseline
prefetch.sectionLimit100Max results rendered per baseline section
json
{
  "limits": {
    "mergedEvidenceCeiling": 1000,
    "prefetch": { "jikjiTopK": 12, "sectionLimit": 30 }
  }
}

The baseline renders each prefetched result's full chunk content (not a per-result excerpt) so the model can see where the hit came from. Bound baseline size with prefetch.sectionLimit / prefetch.minSyncTopK / prefetch.jikjiPathLimit, not with a per-result content truncation.

Ingest caps are trusted connector options under datasources.<name>.connector and are never settable from model/tool arguments: maxDocuments, maxItemsPerFeed, and maxContentChars (RSS); maxDocuments, maxContentChars, maxResultsPerQuery, and maxBytesPerFile (Spotlight); maxDocuments, maxContentChars, maxBytesPerFile, concurrency, bandwidthLimit, and dryRun (cloud-drive/rclone). maxContentChars defaults to 20000 for RSS and 100000 for Spotlight and cloud-drive.

Probe and configure datasource skills (setup wizard)

Use autorag setup to probe the local runtime, model profile, and known datasources automatically:

bash
autorag setup --format json

A datasource setup UI is not shipped in this build — do not recommend it for datasource setup. Configure datasources directly in trusted config, wizard-style:

  1. Probe every datasource for setup feasibility before asking the user anything: the backing CLI exists (lazykatok, discrawl, slacrawl, wacrawl, telecrawl, notcrawl, qmd, mailcrawl, rclone, lark-cli) and its local store or archive is present. CLI-backed datasources own their own archive, index, and authentication, so environment credentials (such as bot tokens) are never required or checked for them. Non-CLI connectors (such as github) require their credential environment variable (GITHUB_TOKEN).
  2. Auto-configure every datasource that probes feasible — write its trusted datasources entries without asking. For example, when Slack (slacrawl) and Discord (discrawl) are installed with local stores present, set both up automatically. Discord uses discrawl's local desktop wiretap archive; no Discord bot token is configured or needed.
  3. Skip every datasource that probes infeasible (for example Notion or Telegram when their CLIs or native stores are not present) and always report the skipped list to the user, with what is missing for each.
  4. Set up a skipped datasource only when the user explicitly asks for it: install or authenticate the backing CLI first, then configure it.
  5. E-mail datasources (mail-export, mailcrawl) matter to most users — always probe them and report their status, even when they end up skipped.

Datasource skills belong in trusted config. Builtin template names are kakao, whatsapp, telegram, slack, discord, clawgallery, notion, github, github-gist, cloud-drive, mail-export, mailcrawl, obsidian, rss, spotlight, and lark. Config keys may be connection aliases with "type": "<template>". Unknown names are skipped with an unknown-datasource-skill warning; they do not fail config resolution. scope narrows a query to a sub-path as ordinary filtering. Tags are descriptive metadata only, not search filters. MCP datasourceIds selects configured connections before retrieval; discover their IDs with autorag.datasources.list.

jsonc
{
  "datasources": {
    "github": { "connector": { "repos": ["owner/repo"], "tokenEnv": "GITHUB_TOKEN" } },
    "github-gist": { "connector": { "tokenEnv": "GITHUB_TOKEN" } },
    "google-drive": { "type": "cloud-drive", "connector": { "provider": "google-drive", "remote": "gdrive:" } },
    "archive-drive": { "type": "cloud-drive", "connector": { "remote": "archive:" } },
    "mailcrawl": { "instanceId": "personal", "connector": { "account": "personal", "mailbox": "INBOX", "binaryPath": "mailcrawl" } },
    "obsidian": { "connector": { "vaultPath": "/path/to/vault" } },
    "rss": { "connector": { "feeds": [{ "url": "https://example.com/feed.xml" }] } }
  }
}

Tokens are environment-variable names, not raw secrets. CLI-backed connectors keep authentication in their external tool configuration.

Mailcrawl must be installed separately (@nomadamas/mailcrawl@0.2.0 or newer) and configured through its own Himalaya account. AutoRAG runs its local sync and index lifecycle, then uses the mailcrawl CLI for BM25, semantic, or hybrid search. Do not use 0.1.3 or earlier: a no-op sync followed by index fails with text array must be non-empty. 0.2.0 defaults to the in-process native Qwen/Qwen3-Embedding-0.6B embedder and keeps vectors in LanceDB, so a cold cache makes the first index download ONNX weights and run for minutes. Use mailcrawl for Gmail, IMAP, and Maildir retrieval. The former Gmail REST datasource is removed.

Verify and build indexes

Configuration alone is not a successful setup:

bash
autorag setup --format json
autorag status --json
autorag health --json
autorag refresh --json
autorag search "summarize the collection" --top-k 3 --json --debug
  • setup probes runtime health, model profile, and datasource readiness.
  • status is model-free and path-opaque.
  • health resolves the single model, checks credential presence, and normally performs one live completion probe.
  • health --skip-probes is only for intentionally offline validation and does not prove live provider access.
  • refresh syncs parsed mirrors, MinSync, Jikji, configured datasources, and on Windows the bundled Everything file-name index. --method <csv> may deliberately narrow it (parsed,minsync,datasources,jikji,everything,all).
  • Use refresh --force for a full resync only when incremental refresh is not enough. Keep destructive reset/rebuild operations scoped to workspace .autorag indexes, never source documents.
  • After a search, use autorag evidence <sessionId> --json to inspect the exact source chunks behind numbered results, including source, method, stable evidence ID, excerpt/content, chunk index, and line number.

Connect external agents through MCP

When setting up AutoRAG for a coding agent, register the package's autorag-mcp stdio executable with the same absolute AUTORAG_CONFIG path. Follow autorag-lite-setup's MCP registration and verification procedure: inspect existing host registration, use an absolute executable path, reload/reconnect, discover schemas with tools/list, and exercise autorag.status, autorag.datasources.list, and a known-phrase autorag.search through MCP. Restart the MCP server after config changes. The MCP server returns model-free source chunks; the calling agent curates them. It does not invoke the configured librarian model. The same server also exposes autorag.report and autorag.evidence for the curation lifecycle; the matching CLI commands remain a maintenance path. Discover the exact schemas with MCP tools/list. Keep the full autorag skill only when model-backed curated search is also wanted. For MCP-only setup, use autorag-lite-setup instead of requiring live model health.

Keep indexes fresh

For continuous freshness, create or verify an OS-appropriate scheduled autorag watch --once job, hourly by default (every 1 hour; shorten only when the user asks for fresher indexes). Prefer cron or launchd on macOS, cron or a user systemd timer on Linux, and Task Scheduler on Windows. Use the same config as search, avoid overlapping runs, and keep logs outside source trees. Once the schedule is installed, tell the user right away that hourly freshness is set up.

Environment overrides

  • AUTORAG_HOME
  • AUTORAG_CONFIG
  • AUTORAG_SEARCH_PATHS
  • AUTORAG_WORKSPACE
  • AUTORAG_MEMORY_PATH
  • AUTORAG_MODEL_PROVIDER
  • AUTORAG_MODEL_ID

Completion condition

Setup is complete only when the CLI is installed, roots are approved, a non-secret single-model config is written, every datasource has been probed and the auto-configured and skipped lists reported to the user, dupey is installed or its absence reported, status is acceptable, live health passes, refresh builds the requested indexes, one real structured search succeeds, and any requested ongoing schedule is installed or verified with the user told it is active. For external-agent integration, also require successful MCP discovery and a known-source MCP search; CLI success alone does not prove the host connection.

© Marker-Inc-Korea, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/autorag-setup of Marker-Inc-Korea/AutoRAG.

Open the folder on GitHubat commit 29ccab9

Compare with similar skills

AutoRAG Setup and Repair next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

AutoRAG Setup and Repair compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
AutoRAG Setup and Repair this skillMarker-Inc-Korea/AutoRAG5.1k—~5.6kAutomated safety check: PassMIT
Sciverseopendatalab/Sciverse-Agent-Tools120—~3kAutomated safety check: PassCustom licence
Neurolink Guidejuspay/neurolink148—~1.4kAutomated safety check: PassMIT
Dive Into LangGraphluochang212/dive-into-langgraph457—~837Automated safety check: NotesCustom licence
MCP Local RAGshinpr/mcp-local-rag412—~4.4kAutomated safety check: PassMIT
Local RAG Searchnkapila6/mcp-local-rag1341 repos~1.6kAutomated safety check: PassMIT

Similar skills

  • Sciverse

    opendatalab/Sciverse-Agent-Tools

    A skill your agent uses when the user needs academic paper retrieval — searching scientific literature by author/year/journal, finding paper chunks for RAG-style citations, or expanding original…

    120 GitHub stars~3k tokensUpdated 21 days ago
    AI & LLM EngineeringAuto-check passed
  • Neurolink Guide

    juspay/neurolink

    Guide for using the NeuroLink SDK and CLI. An agent skill from juspay/neurolink.

    148 GitHub stars~1.4k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Dive Into LangGraph

    luochang212/dive-into-langgraph

    A Chinese-language guide and reference for building agents with LangGraph 1.0, from a first ReAct agent through middleware, memory, MCP, RAG and web search.

    457 GitHub stars~837 tokensUpdated 29 days ago
    AI & LLM EngineeringAuto-check: notes
  • MCP Local RAG

    shinpr/mcp-local-rag

    Searches, saves, and maintains a local document index through a local RAG MCP server.

    412 GitHub stars~4.4k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Local RAG Search

    nkapila6/mcp-local-rag

    Efficiently perform web searches using the mcp-local-rag server with semantic similarity ranking.

    134 GitHub starsUsed in 1 repo~1.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Clawmem

    yoloshii/ClawMem

    ClawMem operational reference for agents at query time — the 3-rule escalation gate, MCP tool routing, the 4 query-optimization levers, pipeline behavior (query vs intentsearch), composite scoring…

    210 GitHub stars~7.5k tokensUpdated today
    Agent WorkflowsAuto-check passed

More from Marker-Inc-Korea/AutoRAG

  • AutoRAG Librarian

    Marker-Inc-Korea/AutoRAG

    Searches, summarizes, compares and answers questions from an already configured AutoRAG librarian agent over local documents and authorized datasources.

    5.1k GitHub stars~2.1k tokensUpdated today
    Auto-check passed
  • AutoRAG Doctor

    Marker-Inc-Korea/AutoRAG

    Diagnoses and repairs a broken AutoRAG install so every configured datasource is both indexed and returns real search hits.

    5.1k GitHub stars~3.4k tokensUpdated today
    Auto-check passed
  • AutoRAG Lite Setup

    Marker-Inc-Korea/AutoRAG

    Bootstraps and repairs the model-free AutoRAG Lite MCP server: installing it, initializing a config with approved search roots, building indexes and verifying discovery.

    5.1k GitHub stars~3.3k tokensUpdated today
    Auto-check passed

Questions about AutoRAG Setup and Repair

What does AutoRAG Setup and Repair do?

Installs, configures, and repairs AutoRAG's search model, approved folders, indexes, and datasources, and registers its Lite MCP server. This skill installs the CLI with Bun or npm if missing, then inspects existing configuration from a config file or environment variable, covering search paths, the chosen model, BM25 and sync settings, retrieval limits, and datasources, and preserves a working config and the user's explicit choices unless they ask to replace it or health checks fail.

When should I use AutoRAG Setup and Repair?

AutoRAG Setup and Repair fits situations like: installing AutoRAG for the first time; fixing a failed init, refresh, or health check; adding a new folder or datasource to AutoRAG's index.

How do I install AutoRAG Setup and Repair in Claude Code?

Run `npx skills add Marker-Inc-Korea/AutoRAG --skill autorag-setup -a claude-code`. Or copy the skill folder (skills/autorag-setup in Marker-Inc-Korea/AutoRAG) into .claude/skills/autorag-setup in your project. Claude Code loads it when a task matches its description.

How do I install AutoRAG Setup and Repair in Codex?

Run `npx skills add Marker-Inc-Korea/AutoRAG --skill autorag-setup -a codex`. Or copy the skill folder (skills/autorag-setup in Marker-Inc-Korea/AutoRAG) into .agents/skills/autorag-setup in your project. Codex loads it when a task matches its description.

Can I use AutoRAG Setup and Repair in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Marker-Inc-Korea/AutoRAG --skill autorag-setup -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/autorag-setup, .gemini/skills/autorag-setup, .github/skills/autorag-setup and .opencode/skills/autorag-setup in your project.

What does AutoRAG Setup and Repair need to run?

Going by SKILL.md and its folder, AutoRAG Setup and Repair needs the command-line tools its instructions call (bun and npm) and credentials named OPENROUTER_API_KEY, GITHUB_TOKEN, TYPESAFE_API_KEY and AI_GATEWAY_API_KEY. Our summary lists: Node.js 24 or newer, or Bun.

Does AutoRAG Setup and Repair access the network?

SKILL.md names 1 domain. In commands or code: openrouter.ai; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is AutoRAG Setup and Repair safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does AutoRAG Setup and Repair use?

AutoRAG Setup and Repair is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does AutoRAG Setup and Repair use?

About 5.6k tokens (SKILL.md is roughly 22k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to AutoRAG Setup and Repair?

Skills that share tags, products or a category with AutoRAG Setup and Repair: Sciverse (opendatalab/Sciverse-Agent-Tools, 120 stars), Neurolink Guide (juspay/neurolink, 148 stars), Dive Into LangGraph (luochang212/dive-into-langgraph, 457 stars) and MCP Local RAG (shinpr/mcp-local-rag, 412 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains AutoRAG Setup and Repair?

Marker-Inc-Korea (a GitHub organization) maintains it in Marker-Inc-Korea/AutoRAG, which has 5,122 GitHub stars. The repository holds 4 skills in this directory. The repository was last updated on October 11, 2026.

Source: Marker-Inc-Korea/AutoRAG on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.