Agent skill

AutoRAG Doctor

by Marker-Inc-Korea in Marker-Inc-Korea/AutoRAG

Diagnoses and repairs a broken AutoRAG install so every configured datasource is both indexed and returns real search hits.

MITAuto-check passedAI & LLM Engineering

Install AutoRAG Doctor

skills CLI
$ npx skills add Marker-Inc-Korea/AutoRAG --skill autorag-doctor -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Marker-Inc-Korea/AutoRAG autorag-doctor --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Marker-Inc-Korea/AutoRAG.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/autorag-doctor .claude/skills/autorag-doctor && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
autorag-doctor
GitHub stars
5.1k
Token cost
~3.4k tokens
SKILL.md length
1,625 words
Files
1
Skills in repo
4
Repo updated
First seen
Licence
MIT

At a glance

Diagnoses and repairs a broken AutoRAG install so every configured datasource is both indexed and returns real search hits.

  • Works in 4 steps: Triage the core → Probe every datasource natively → Prove searchability → …
  • Search returns nothing or too little
  • SKILL.md covers Safety, 1. Triage the core, 2. Probe every datasource… and 3. Prove searchability, plus 4 more sections
  • Calls cargo; needs OPENROUTER_API_KEY

What it does

The goal is that each datasource is indexed and actually answers a query; a green status alone is not enough. The agent triages the core with status, health and gateway status commands in JSON, reads the config for search paths, workspace, MinSync, Jikji and datasources, then probes every datasource with its own native check, since each CLI owns its archive. If the CLI or config is missing it stops and points to the setup skill, because the doctor repairs an install and does not create one.

Repairs cover orphan locks and processes, embedding dimension or identity mismatches, stale indexes and missing setup, with re-sync through refresh by method. Safety rules forbid deleting or editing source documents or a datasource's native store, allow resetting only AutoRAG-owned state, never print token values and kill a process only after confirming it is an AutoRAG orphan. Datasources include CLIs such as qmd, rclone and Spotlight, and the run always ends with a status table.

When your agent uses it

  • Search returns nothing or too little
  • A refresh hangs or fails
  • A datasource has disappeared from the results
  • The embedding gateway will not start
  • Indexes look stale after the corpus changed

Example prompts

  • “Check my AutoRAG install and make sure every datasource returns hits.”
  • “Refresh keeps hanging, so find any orphan locks or processes.”
  • “Search no longer shows my Slack archive; diagnose and fix it.”

Requirements

  • The autorag CLI and an existing config
  • The CLIs of the datasources you configured

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Triage the core
  2. Probe every datasource natively
  3. Prove searchability
  4. Repair playbook

What it can do on your machine

Read from SKILL.md and the folder at commit 29ccab9. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • cargo

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • OPENROUTER_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

AutoRAG Doctor loads about 3.4k tokens when it runs. Until then it costs about 166 tokens; SKILL.md has 1,625 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~166
When it runs · the whole SKILL.md, loaded when a task matches
~3.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Marker-Inc-Korea/AutoRAG at commit 29ccab9, republished under its MIT licence (© Marker-Inc-Korea). 1,625 words, ~3,437 tokens.

Download SKILL.mdSave it as .claude/skills/autorag-doctor/SKILL.md (or your agent's skills folder).
name
autorag-doctor
description
Diagnose and repair a broken or half-working AutoRAG install so every configured source is both indexed and searchable. Checks AutoRAG, MinSync, Jikji, the embedding gateway, and every CLI-backed datasource (lazykatok, discrawl, slacrawl, wacrawl, telecrawl, notcrawl, qmd, mailcrawl, rclone, Spotlight), then fixes orphan locks, orphan processes, embedding dimension or identity mismatches, stale indexes, and missing setup. Use when search returns nothing or too little, refresh hangs or fails, a datasource disappeared from results, indexes look stale, the gateway will not start, or the user asks to check, diagnose, verify, or repair AutoRAG.
license
MIT

AutoRAG doctor

Goal: every configured datasource is indexed and actually returns hits. A green status is not enough — the run is done only when a real query returns a real hit per datasource, or the datasource is reported as genuinely empty or not configured.

Always finish with the status table in Report.

Safety

  • Never delete, move, or edit source documents.
  • Never delete a datasource's native store (~/Library/Application Support/katok, ~/.discrawl, ~/.mailcrawl, .qmd, Telegram/WhatsApp/Notion archives). AutoRAG only reads them; rebuilding them is the owning CLI's job.
  • Only AutoRAG-owned state under AUTORAG_HOME and the workspace .autorag directory may be reset.
  • Never print token or password values. Report credential names only.
  • Kill a process only after confirming it is an AutoRAG-owned orphan.

1. Triage the core

bash
autorag status --json
autorag health --json
autorag gateway status --format json
  • status reports state, stale, diagnostics, and per-component state for minsync, jikji, and datasources. stale: true or any stale-index diagnostic means the corpus changed since the last successful refresh.
  • health resolves the single search model and does one live completion probe. --skip-probes only proves config shape, never live access; do not claim a healthy model from it.
  • The gateway is on-demand: stopped is normal when nothing is embedding. unavailable while a refresh is running is a real failure.

Config lives at --config, AUTORAG_CONFIG, $AUTORAG_HOME/config.json, or ~/.autorag/config.json. Read searchPaths, workspacePath, minSync, jikji, and datasources before changing anything.

If the CLI itself is missing or the config does not exist, stop and run the autorag-setup skill first — doctor repairs an existing install, it does not create one.

2. Probe every datasource natively

Each CLI owns its archive, so ask the CLI, not AutoRAG. A datasource is only active when its binary exists, its store is present, and its own check passes.

DatasourceNative checkRe-sync when empty or stale
MinSync (local docs)minsync status, minsync check, minsync verifyautorag refresh --method minsync
Jikji (discovery)jikji doctorautorag refresh --method jikji
Everything (Windows file names)autorag status --json → components.everythingautorag refresh --method everything --json
KakaoTalklazykatok doctorlazykatok sync && lazykatok index
Discorddiscrawl --json metadatadiscrawl sync
Slackslacrawl --json doctorslacrawl sync
WhatsAppwacrawl --json doctorwacrawl import
Telegramtelecrawl --json doctor, telecrawl --json statustelecrawl import
Notionnotcrawl doctor, notcrawl statusnotcrawl sync --source desktop
Obsidian / notesqmd statusqmd update && qmd embed
Mailmailcrawl doctor, mailcrawl statusmailcrawl sync && mailcrawl index
Cloud driverclone listremotesautorag refresh --method datasources
Spotlight (macOS)mdutil -s /indexed by the OS; no AutoRAG sync
Lark / Feishulark-cli auth status --format jsonno local sync; search is remote

Rules:

  • A missing binary is not configured, not a failure. Report it with the install command and move on.
  • An empty store is a legitimate zero-hit result. telecrawl and slacrawl return JSON null (not []) for no hits — treat that as empty, not broken.
  • A configured datasource whose native check fails is a failure and must be repaired or reported explicitly. Never report a skip as a pass.
  • Zero hits with a non-empty store means the CLI's search path is suspect, not the data. Cross-check against the raw index before concluding (for example slacrawl sql "select count(*) from message_fts where message_fts match 'x';" against slacrawl search x). A populated index plus an empty CLI result is an upstream bug: report it with the exact reproduction instead of re-syncing.

3. Prove searchability

Indexing without retrieval is a failed run. Probe retrieval per datasource through the model-free MCP tools, then once end to end:

text
autorag.status {}
autorag.search {"query":"a word that certainly appears","topK":3}
autorag.search {"query":"recent topic","datasourceIds":["discord"],"topK":3}
autorag.search {"query":"recent mail subject","scope":"/mailcrawl/**","topK":3}
bash
autorag search "summarize the collection" --top-k 3 --json --debug
  • Always read the diagnostics returned by MCP autorag.search. A run that silently dropped a whole retrieval method still looks successful, just with fewer results; only the diagnostics name it (retrieval-method-failed). The model-backed CLI autorag search --json hides diagnostics, sessionId, and per-result evidence unless --debug is set, so pass --debug when diagnosing that path.
  • A method missing from the returned method values means that method contributed nothing. During a full MinSync re-sync this is expected: the store is being rebuilt, minsync status reports NotSynced, and local-file hits stay absent until it finishes. Confirm with minsync status before treating it as a failure, and never kill a running sync to "fix" it.
  • MCP autorag.search needs no model, so it isolates retrieval from model failures.
  • Use MCP datasourceIds to select configured connections before retrieval; scope narrows results within scope-capable datasources. Discover connection IDs with autorag.datasources.list; descriptor tags are metadata only, not search filters. Every configured connection is searchable — if one returns nothing, investigate its native store, connector, or the query itself.
  • autorag.evidence {"sessionId":"...","resultNumber":N} shows the exact chunk behind a numbered result; use it to confirm a hit is real and its source is readable. The CLI autorag evidence SESSION --json remains for terminal repair.
  • Local-file hits must map to an absolute, existing path. Datasource hits keep source-native identities such as /kakao/personal/chunks/42; those are not filesystem paths and must never be passed to cat.

4. Repair playbook

Orphan lock

The embedding runtime keeps embedding-runtime.lock, embedding-runtime.pid, and embedding-runtime.port in AUTORAG_HOME. It reclaims them automatically when the recorded PID is dead; a lock-conflict means the PID is alive.

bash
autorag gateway status --format json
autorag gateway stop

Only when gateway stop cannot clear it, and the recorded PID is confirmed dead (kill -0 PID fails), remove the three files by hand and retry.

Index locks are owned by their engines: MinSync/tantivy locks under <workspace>/.autorag/, and per-CLI locks such as <workspace>/.autorag/datasources/discrawl/.discrawl-sync.lock. Delete one only after confirming no owning process is alive; otherwise wait for the run that holds it.

Orphan process
bash
pgrep -fl 'autorag|autorag-gateway|minsync' | grep -v pgrep

A refresh that was killed mid-run can leave the gateway or a minsync child alive. Stop the gateway with autorag gateway stop first; only kill a PID directly when it is confirmed orphaned (no parent CLI, no live refresh). Re-run autorag status --json afterwards to confirm inFlight: false.

Show full SKILL.md (687 more words)Show less
Embedding model mismatch

embedding-identity-mismatch or a dimension error means the vectors on disk were built with a different embedder than the configured one. Vectors of two different dimensions can never be compared, so the index must be rebuilt:

bash
autorag models prefetch --profile qwen3-embedding-0.6b
autorag index rebuild --method minsync

The current default is the local native:Qwen/Qwen3-Embedding-0.6B runtime at 1024 dimensions. A workspace still pinned to the legacy 768-dimension Ollama/TEI path must be reindexed explicitly, or pinned to an explicit profile. Changing embedder.dimension in config without a rebuild leaves retrieval silently empty. Keep embeddings local; do not point the embedder at a remote endpoint to work around a local failure.

Stale index

stale-index diagnostics or stale: true mean sources changed after the last refresh.

bash
autorag refresh --method parsed,minsync --json
autorag refresh --force --json

Use --force only when incremental refresh does not clear it. For continuous freshness install an hourly autorag watch --once job (cron, launchd, systemd timer, or Task Scheduler).

Missing or wrong setup
  • unknown-datasource-skill: the config key is not a builtin template and has no "type". It is skipped, not fatal — fix the name or add "type".
  • datasource-index-failed / sync-failed: the backing CLI errored. Run that CLI's own check from the table above and fix it there.
  • A missing MinSync binary is not a diagnostic: autorag search, autorag refresh and MCP tools fail with MinSync is required ... and a non-zero exit. Install it (cargo install minsync) or leave minSync.autoInstall on, then retry.
  • jikji-unavailable: the binary is missing and auto-install failed. It installs through cargo; verify the Rust toolchain, then autorag refresh --method jikji to retry.
  • auth-error / rate-limited: model or datasource credentials. Report the missing environment-variable name and let the user supply it.
  • A configured datasource that returns nothing is a native store, connector, or query problem — every configured connection is searchable. Run its native check from the table above and fix it there.

Diagnostic codes

CodeMeaningFirst move
stale-indexSources changed since last refreshautorag refresh --method parsed,minsync
index-not-readyIndex missing or never builtautorag refresh --json
minsync-sync-failedMinSync indexing failed; the message carries MinSync's own reasonFix the reported cause, then autorag refresh --method minsync
jikji-unavailableJikji binary missing or install failedCheck cargo, retry refresh
everything-index-failedWindows Everything instance could not start or index; message carries ES exit code and stderrFix the reported cause, autorag refresh --method everything --json
embedding-identity-mismatchIndexed vectors use a different embedderautorag index rebuild --method minsync
embedder-unavailableEmbedding gateway or runtime downautorag gateway status --format json
lock-conflictAnother runtime holds the lockautorag gateway stop
datasource-index-failedBacking CLI failed to indexRun that CLI's own doctor
datasource-emptyStore has no matching contentRe-sync with the CLI, or accept as empty
unknown-datasource-skillConfig name is not a known templateFix the name or add "type"
retrieval-method-failedOne method errored during the queryRead --debug diagnostics
auth-errorCredentials missing or rejectedReport the env var name
query-route-fallbackJev routing unavailable (often OPENROUTER_API_KEY unset); searched local with the original questiontest -n "$OPENROUTER_API_KEY"; report the env var name, never its value
query-decomposition-failedDecomposition model call failed; searched the original questionCheck queryDecomposition.model resolves (autorag models list --provider openrouter)
follow-up-check-fallbackJev post-fast-answer check unavailable; the run verifiedSame as query-route-fallback
datasource-selection-fallbackJev datasource check unavailable; no datasource was searched before the fast answerSame as query-route-fallback
query-routed / datasources-selected / follow-up-skippedInfo: Jev's branch and queries / datasources searched and skipped / fast answer judged finalNone; working as intended. A datasource that is never selected usually needs a clearer description in the config

Report

Always end with this table, one row per datasource and per AutoRAG component:

SourceConfiguredIndexedSearchableIssue foundFix applied
minsyncyesyesyes (3 hits)––
discordyesyesnostale archivediscrawl sync
telegramno––CLI not installedreported

Searchable must come from an actual query in step 3, never inferred from index state. Follow the table with the exact remaining action for every row that is not fully green.

Completion condition

Done only when: core triage is clean or every remaining diagnostic is explained, every configured datasource passed its native check, every one of them returned a real hit or is proven empty, every repair was re-verified by re-running the failing check, and the report table was delivered.

© Marker-Inc-Korea, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/autorag-doctor of Marker-Inc-Korea/AutoRAG.

Open the folder on GitHubat commit 29ccab9

Compare with similar skills

AutoRAG Doctor next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

AutoRAG Doctor compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
AutoRAG Doctor this skillMarker-Inc-Korea/AutoRAG5.1k—~3.4kAutomated safety check: PassMIT
Chroma Vector DatabaseOrchestra-Research/AI-Research-SKILLs13k7 repos~2.3kAutomated safety check: PassMIT
Embeddings via 9Routerdecolua/9router31k—~604Automated safety check: PassMIT
AI SDK Developmenttrypostit/trypost6921 repos~3.5kAutomated safety check: PassMIT
Retail Product Search Agentgoogle/adk-recipes10k—~3kAutomated safety check: PassApache-2.0
Evaluate RAGai-evals-course/evals-skills1.5k—~1.9kAutomated safety check: PassApache-2.0

Similar skills

  • Chroma Vector Database

    Orchestra-Research/AI-Research-SKILLs

    Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.

    13k GitHub starsUsed in 7 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Embeddings via 9Router

    decolua/9router

    Generates vector embeddings through the 9Router /v1/embeddings endpoint, using models from providers such as OpenAI, Gemini, Mistral and Voyage for RAG and semantic search.

    31k GitHub stars~604 tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check passed
  • AI SDK Development

    trypostit/trypost

    TRIGGER when working with ai-sdk which is Laravel official first-party AI SDK.

    692 GitHub starsUsed in 1 repo~3.5k tokens
    AI & LLM EngineeringAuto-check passed
  • Retail Product Search Agent

    google/adk-recipes

    Official

    Builds a retail product search agent on Google Cloud, from catalog ingestion into BigQuery and Vector Search to ADK scaffolding, evaluation and Cloud Run deployment.

    10k GitHub stars~3k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Evaluate RAG

    ai-evals-course/evals-skills

    Guides evaluation of a RAG system by diagnosing failures in traces, building a retrieval test set and scoring retrieval and generation separately.

    1.5k GitHub stars~1.9k tokensUpdated 16 days ago
    AI & LLM EngineeringAuto-check passed
  • Michel CLI Demo Recorder

    PackmindHub/packmind

    Produce proof-of-execution demos of the Packmind CLI (packmind-cli) as terminal-styled images (colors and formatting preserved exactly), for embedding in a GitHub PR.

    318 GitHub stars~3.4k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed

More from Marker-Inc-Korea/AutoRAG

  • AutoRAG Librarian

    Marker-Inc-Korea/AutoRAG

    Searches, summarizes, compares and answers questions from an already configured AutoRAG librarian agent over local documents and authorized datasources.

    5.1k GitHub stars~2.1k tokensUpdated today
    Auto-check passed
  • AutoRAG Lite Setup

    Marker-Inc-Korea/AutoRAG

    Bootstraps and repairs the model-free AutoRAG Lite MCP server: installing it, initializing a config with approved search roots, building indexes and verifying discovery.

    5.1k GitHub stars~3.3k tokensUpdated today
    Auto-check passed
  • AutoRAG Setup and Repair

    Marker-Inc-Korea/AutoRAG

    Installs, configures, and repairs AutoRAG's search model, approved folders, indexes, and datasources, and registers its Lite MCP server.

    5.1k GitHub stars~5.6k tokensUpdated today
    Auto-check passed

Questions about AutoRAG Doctor

What does AutoRAG Doctor do?

Diagnoses and repairs a broken AutoRAG install so every configured datasource is both indexed and returns real search hits. The goal is that each datasource is indexed and actually answers a query; a green status alone is not enough. The agent triages the core with status, health and gateway status commands in JSON, reads the config for search paths, workspace, MinSync, Jikji and datasources, then probes every datasource with its own native check, since each CLI owns its archive.

When should I use AutoRAG Doctor?

AutoRAG Doctor fits situations like: search returns nothing or too little; A refresh hangs or fails; A datasource has disappeared from the results; the embedding gateway will not start.

How do I install AutoRAG Doctor in Claude Code?

Run `npx skills add Marker-Inc-Korea/AutoRAG --skill autorag-doctor -a claude-code`. Or copy the skill folder (skills/autorag-doctor in Marker-Inc-Korea/AutoRAG) into .claude/skills/autorag-doctor in your project. Claude Code loads it when a task matches its description.

How do I install AutoRAG Doctor in Codex?

Run `npx skills add Marker-Inc-Korea/AutoRAG --skill autorag-doctor -a codex`. Or copy the skill folder (skills/autorag-doctor in Marker-Inc-Korea/AutoRAG) into .agents/skills/autorag-doctor in your project. Codex loads it when a task matches its description.

Can I use AutoRAG Doctor in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Marker-Inc-Korea/AutoRAG --skill autorag-doctor -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/autorag-doctor, .gemini/skills/autorag-doctor, .github/skills/autorag-doctor and .opencode/skills/autorag-doctor in your project.

What does AutoRAG Doctor need to run?

Going by SKILL.md and its folder, AutoRAG Doctor needs the command-line tools its instructions call (cargo) and credentials named OPENROUTER_API_KEY. Our summary lists: The autorag CLI and an existing config; The CLIs of the datasources you configured.

Does AutoRAG Doctor access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is AutoRAG Doctor safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does AutoRAG Doctor use?

AutoRAG Doctor is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does AutoRAG Doctor use?

About 3.4k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to AutoRAG Doctor?

Skills that share tags, products or a category with AutoRAG Doctor: Chroma Vector Database (Orchestra-Research/AI-Research-SKILLs, 13k stars), Embeddings via 9Router (decolua/9router, 31k stars), AI SDK Development (trypostit/trypost, 692 stars) and Retail Product Search Agent (google/adk-recipes, 10k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains AutoRAG Doctor?

Marker-Inc-Korea (a GitHub organization) maintains it in Marker-Inc-Korea/AutoRAG, which has 5,122 GitHub stars. The repository holds 4 skills in this directory. The repository was last updated on October 11, 2026.

Source: Marker-Inc-Korea/AutoRAG on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.