Multi-Source to NotebookLM Processor
joeseesun/qiaomu-anything-to-notebooklm
Collects content from WeChat articles, web pages, YouTube, podcasts, documents and more, uploads it to NotebookLM and generates podcasts, slides or mind maps.
Build, extend, AND query a persistent LLM-maintained wiki for any research topic.
$ npx skills add iusztinpaul/ai-research-os-workshop --skill research -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install iusztinpaul/ai-research-os-workshop research --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/iusztinpaul/ai-research-os-workshop.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/ai-research-os/skills/research .claude/skills/research && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "research" agent skill from https://github.com/iusztinpaul/ai-research-os-workshop/tree/main/plugins/ai-research-os/skills/research into .claude/skills/research/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "research", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/iusztinpaul/ai-research-os-workshop/tree/main/plugins/ai-research-os/skills/researchType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add iusztinpaul/ai-research-os-workshop --skill research -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install iusztinpaul/ai-research-os-workshop research --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/iusztinpaul/ai-research-os-workshop.git skills-src && mkdir -p .agents/skills && cp -r skills-src/plugins/ai-research-os/skills/research .agents/skills/research && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "research" agent skill from https://github.com/iusztinpaul/ai-research-os-workshop/tree/main/plugins/ai-research-os/skills/research into .agents/skills/research/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "research", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add iusztinpaul/ai-research-os-workshop --skill research -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install iusztinpaul/ai-research-os-workshop research --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/iusztinpaul/ai-research-os-workshop.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/plugins/ai-research-os/skills/research .cursor/skills/research && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "research" agent skill from https://github.com/iusztinpaul/ai-research-os-workshop/tree/main/plugins/ai-research-os/skills/research into .cursor/skills/research/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "research", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/iusztinpaul/ai-research-os-workshop.git --path plugins/ai-research-os/skills/research--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add iusztinpaul/ai-research-os-workshop --skill research -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install iusztinpaul/ai-research-os-workshop research --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/iusztinpaul/ai-research-os-workshop.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/plugins/ai-research-os/skills/research .gemini/skills/research && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "research" agent skill from https://github.com/iusztinpaul/ai-research-os-workshop/tree/main/plugins/ai-research-os/skills/research into .gemini/skills/research/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "research", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install iusztinpaul/ai-research-os-workshop researchInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add iusztinpaul/ai-research-os-workshop --skill research -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/iusztinpaul/ai-research-os-workshop.git skills-src && mkdir -p .github/skills && cp -r skills-src/plugins/ai-research-os/skills/research .github/skills/research && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "research" agent skill from https://github.com/iusztinpaul/ai-research-os-workshop/tree/main/plugins/ai-research-os/skills/research into .github/skills/research/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "research", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add iusztinpaul/ai-research-os-workshop --skill research -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install iusztinpaul/ai-research-os-workshop research --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/iusztinpaul/ai-research-os-workshop.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/plugins/ai-research-os/skills/research .opencode/skills/research && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "research" agent skill from https://github.com/iusztinpaul/ai-research-os-workshop/tree/main/plugins/ai-research-os/skills/research into .opencode/skills/research/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "research", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
researchBuild, extend, AND query a persistent LLM-maintained wiki for any research topic.
Research is an agent skill from iusztinpaul/ai-research-os-workshop. Build, extend, AND query a persistent LLM-maintained wiki for any research topic. Conversational entry point that routes between four modes - query (fast read-only answer from existing wiki), append (ingest the sources the user provided, no discovery), deep (explicit discovery/deep research at a fast/light/deep depth preset), and init (create a new research directory). Ingests from your knowledge sources (Obsidian vault + Readwise + NotebookLM + GitHub repos + YouTube videos + web seeds + user-dropped PDFs) and…
Its SKILL.md is about 17k tokens, which your agent loads only when the skill is triggered. The skill folder holds 19 other files, including scripts (for example `CONVENTIONS.md`, `agents/builder.md` and `agents/comparison_writer.md`).
It sits in Knowledge Management, covering Deep research, Source-grounded notebooks and PDF. It works with NotebookLM, GitHub, Obsidian and YouTube. The repository describes itself as: How to turn your Second Brain into a living research memory that your agents maintain. Workshop with slides, video and code. The licence is MIT.
12 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit dc66605. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 8 files in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
uvjqcurlpython3npmFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
github.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Research loads about 17k tokens when it runs. Until then it costs about 246 tokens; SKILL.md has 8,320 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from iusztinpaul/ai-research-os-workshop at commit dc66605, republished under its MIT licence (© iusztinpaul). 8,320 words, ~17,338 tokens.
.claude/skills/research/SKILL.md (or your agent's skills folder). This skill also uses 17 other files; get the full folder from GitHub.You are a research orchestrator. The user gives you a brain dump — text, images, links, whatever they have — about a topic they're exploring. Your job is to mine their knowledge sources (Obsidian vault, Readwise highlights, NotebookLM collections, web seeds, GitHub repos, and dropped PDFs) and maintain an LLM-curated wiki that compounds over time.
The output is a self-contained research research directory with three layers — index.yaml / index.md (catalog), wiki/ (synthesis), and raw/ (immutable sources). Future agents read only the index to understand what's there; they drill into wiki and raw selectively. The full data contract lives in CONVENTIONS.md.
Before touching source CLIs or writing files, classify the user's intent. Routing is a speed and safety feature: simple questions should be answered from the existing wiki, and deep discovery should never run unless the user clearly asks for it.
| Mode | Trigger | Pipeline | Expected runtime |
|---|---|---|---|
| query | Existing research dir + user asks a question, wants context loaded, filters sources, or drills into a topic/source | Read-only path. See Query path below. Optional Q&A save-back. No source CLI preflight, no discovery, no raw/wiki rewrite. | Seconds to <1 min |
| append | User provides one or more sources to add (drops files/links/repos/videos/PDFs, or says "add this", "just ingest these", "don't run deep research") | Ingest the provided sources only — no discovery, no NLM sweep, no naming sub-tiers. Seeds get relevance_score: 1.0; dedup against index.yaml; Step 1 -> Steps 6-8. | ~1-10 min |
| deep | User explicitly asks for discovery ("deep research", "find more sources", "discover", "exhaustive"), or confirms it at the deep-research gate below. Runs at a depth preset: fast / light / deep. | Discovery. Step 1, Step 2 (pick the preset), Step 3, Step 3b, Step 4, Step 5, then Steps 6-8. Dedup against existing index.yaml. | fast ~5-10 min · light ~10-20 min · deep ~20-40+ min |
| init | No matching research dir exists and the user wants a new research topic | Create working-dir/research-<topic-slug>/, then run append or deep by the same rules (the deep-research gate applies if sources/links were provided). | ~1-10 min append · ~10-40+ min deep |
How to decide:
topic_slug from the user's words (kebab-case).working-dir/research-*/index.yaml for a semantic topic match.Whenever the input carries sources/links to add and the user hasn't already said which way they want it, ask one short question before ingesting anything:
You gave me N source(s). Want me to just ingest them, or run deep research to discover more from your knowledge sources? If deep research, which depth — fast (1 round, 3 queries), light (2 rounds, 3 + 2), or deep (3 rounds, 3 each)?
fast / light / deep) → deep at that preset; you already have the
preset, so skip the Step 2 question.Skip this gate only when the user already made the call explicitly in their message, or for query. This gate is the one place deep research is opt-in — never start discovery rounds without the user choosing them here (or in their original message).
Before any ingest mode writes files or starts expensive work, show a short plan. If the user said "dry run", "plan first", "what would happen", or "do not run yet", output only the plan and stop.
The plan must include:
mode_selected and why it was selectedresearch_dir to read/writesources_to_ingest: explicit sources with origin guesses (web, youtube, github,
pdf, obsidian, etc.) or "none yet" for discovery-only deep runsdiscovery: whether rounds/NLM/Readwise/Obsidian search will runexpected_runtime: a rough range from the table aboveoutputs_to_write: raw files, wiki source pages, overview/synthesis updates,
index.yaml, index.md, log.mdskip_policy: sources already in index.yaml are skipped; unavailable CLIs degrade
with warningsAsk for confirmation before proceeding when any of these are true:
For append, show the plan and proceed unless the user asked for confirmation first. For query, do not show a plan; answer fast.
raw/ and wiki/ directly under the research dir, no memory/ wrapper). If it's older (v1 with raw at root, or v3 with a memory/ wrapper), migrate it to v4 first - or instruct the user to - before ingesting.index.yaml for append and deep modes. If index.yaml is missing but index.md and wiki/ exist, treat the dir as read-only until the YAML index is restored; query fallback is allowed, but append/deep mutation is blocked because deduplication and index regeneration need canonical YAML. The original_path set is the dedup key - sources already there are skipped during research rounds and refused/skipped in append mode with an "already ingested" message.created timestamp for append and deep modes; pass it as --existing-created to build_index_yaml.py in Step 6.7. On these modes also pass the prior index via --existing-index so its existing sources are merged forward (Step 6.7 handles this) — otherwise the rebuild would keep only the newly-ingested sources.In append and seed-only init, skip Steps 2 (configure depth), 3 (initial queries),
3b (NLM discovery), 4 (research rounds), and 5 (merge/dedup/score). The user's seed sources
get relevance_score: 1.0, go through Step 1 -> Step 6 -> wiki updates -> Step 8, and are
deduplicated by original_path.
When intent is genuinely ambiguous, default to query, not ingest. Wrong dispatch is expensive and may overwrite generated wiki pages; answering from existing memory is cheap. Ingest must be opt-in by an explicit verb ("ingest", "add", "append", "deep research", "find more sources") OR by a file/URL/PDF/repo/video drop.
Used when Step 0 selected query mode. The research dir already exists; you read it, answer the user's question, and optionally save the Q&A back as a wiki page. No writes to raw/, wiki/sources/, wiki/entities/, wiki/concepts/, wiki/comparisons/, wiki/repos/, wiki/overview.md, wiki/synthesis.md, or index.yaml ever happen here. Allowed writes: wiki/questions/YYYY-MM-DD-<slug>.md (Q&A save-back), wiki/open-questions.md (when the user explicitly flags an unresolved question), index.md (regenerated), log.md (append).
Use the same locator logic as Step 0 (path provided / topic slug / scan working-dir / ask if multiple).
Read <research_dir>/index.yaml. The YAML is canonical and has the full schema. The MD is for humans browsing in Obsidian.
If index.yaml is missing but index.md and wiki/ exist, continue in read-only fallback:
index.md for topic, source list, scores, origins, and raw/source-page links.wiki/overview.md / wiki/synthesis.md for synthesized answers.index.md, append to log.md, or mutate wiki files in fallback mode.index.yaml is missing, so I answered from index.md + wiki fallback. Restore/regenerate index.yaml before append/deep ingest."Parse and understand:
topic, input_summary, total_sources, total_wiki_pagessources[] — title, origin, original_path, source_url, authors, published_date, publication, relevance_score, summary, tags, uri_full, uri_highlights, uri_source_page, assets, plus origin-specific fields (readwise_location, nlm_*, github_*)For any source, read in this order and stop when the user's question is answered:
summary in index.yaml. Always available, ~2–3 sentences. The default — never go deeper unless there's a reason.uri_source_page (wiki/sources/<slug>.md). LLM-extended summary that's denser than Layer 1 but lighter than Layer 3. Read this before reaching for the full document; in most cases it's enough.uri_highlights. Optional. Exists only when the source carries manually user-curated highlights (typically Readwise-synced). Never LLM-extracted.uri_full. The complete document. Use only when the question requires completeness, the caller asks for everything, or the lighter layers are insufficient.Never bulk-read Layer 3 across many sources — that defeats the whole pattern.
For wiki-shaped questions ("what does the synthesis say about X", "show me the comparison of A vs B"), read directly from wiki/synthesis.md, wiki/overview.md, wiki/comparisons/..., wiki/entities/..., wiki/concepts/..., wiki/contradictions.md, wiki/open-questions.md — these are short by design.
Match the user's request to one of these shapes:
Query mode ("find sources about X"): Scan summaries + tags + entity/concept frontmatter. Return matches at Layer 1.
Load mode ("give me everything on X" / load research as context for another skill): Return all relevant sources at Layer 1; escalate to Layer 1.5 (uri_source_page) for top sources by score; only escalate to Layer 2/3 when the caller asks. Be a clear citizen of the depth knob:
summary → Layer 1 onlywiki → Layer 1 + 1.5 (read source pages for top N)highlights → Layer 1 + 1.5 + Layer 2 (where present), falls through to Layer 3 when Layer 2 is nullfull → escalate to Layer 3 for sources the caller specifies (NOT all sources by default)Filter mode: filter sources by origin, readwise_location, relevance_score >= threshold, tags, seeds-only (relevance_score == 1.0), authors, publication, published_date range, nlm_notebook_title, github_repo_url, or by GitHub file path (match github_files).
Drill-down mode ("tell me more about source X" / "what does the wiki say about concept Y"): For sources, escalate Layer 1 → 1.5 → 2 → 3 only as needed. For wiki concepts/entities, read the page directly.
Compose mode ("summarize what we know about X"): synthesize from the wiki layer (overview / synthesis / relevant entity-concept pages) — these are already the synthesis. If the wiki layer doesn't cover the question, fall back to Layer 1 of relevant sources, then Layer 1.5 if needed.
For GitHub sources: uri_full points at <repo>/ARCHITECTURE.md (a wiki hub). To drill into a specific module, read ARCHITECTURE first, find the inline link to the module doc (e.g., [vectordb](./vectordb.md)), and read it. Module docs are not separately indexed.
After answering, decide whether to save the Q&A as wiki content. Save when EITHER condition is met:
Otherwise, do not save — most questions are conversational and shouldn't compound.
When you save, never put the answer body inside wiki/questions/. The actual knowledge — diagrams, claims, source citations, code permalinks — lands in the wiki at the most appropriate existing location (the knowledge doc). The wiki/questions/ entry stays a slim pointer: the verbatim question, a 1-line why this matters, and a wikilink to the knowledge doc.
This split has two purposes:
| Question scope | Knowledge doc lands at | Notes |
|---|---|---|
| Repo-scoped — drills into a single GitHub source already in the wiki | wiki/repos/<repo>/<TOPIC>.md (e.g. wiki/repos/claude-code/TOOL_PATTERNS.md) | Add a "Deep dive" cross-link from the relevant ARCHITECTURE.md section so a reader scrolling the architecture finds it naturally. |
| Concept / entity drill-down — about a concept or entity the wiki tracks | Enrich existing wiki/concepts/<slug>.md / wiki/entities/<slug>.md in place | If the concept doesn't yet have a page but the answer has enough material to start one, create it. If it's only a single-source mention, flag in wiki/open-questions.md for next ingest instead. |
| Comparison — compares ≥ 2 concepts/entities the wiki tracks | wiki/comparisons/<a-vs-b>.md (create or enrich) | |
| Cross-cutting synthesis — doesn't fit the buckets above | wiki/notes/<topic-slug>.md (create the wiki/notes/ dir if absent) | Catch-all for question-driven knowledge that synthesizes across sources without being scoped to a specific repo / concept / comparison. |
Idempotency rule: if a knowledge doc on the same topic already exists, update it in place. Don't write a new doc just because the question came back. Update its frontmatter (last_updated, append to spawned_by_question), enrich the body, and the new slim question page in wiki/questions/ points at the (now-enriched) doc.
The knowledge doc carries the substance of the answer:
type, name, created, last_updated, and spawned_by_question (path to the slim question page; if multiple questions have enriched this doc, list them).agents/github_spec_writer.md § "Mermaid guidance" (flowchart / sequenceDiagram / classDiagram / mindmap / stateDiagram-v2). Prefer a diagram over a prose paragraph whenever the explanation is structural.[[wiki/sources/<slug>]]) or its raw doc with a heading anchor ([[wiki/repos/<repo>/ARCHITECTURE.md#<heading>|cite]]). Code snippets get commit-pinned permalinks where applicable.> Synthesis: line with your meta-judgment + a one-sentence hint at what new sources would extend or revise this doc.Write to <research_dir>/wiki/questions/YYYY-MM-DD-<question-slug>.md. Slugify the question to ≤ 60 chars (drop articles, lowercase, kebab-case).
---
type: question
name: <verbatim user question, no editorializing>
asked_on: <ISO-8601 date>
sources_cited: [<wiki page paths cited by the answer>]
answer_doc: <wiki path to the knowledge doc>
---
# <verbatim user question>
> Asked on <date>. Answered using <N> source(s) and <M> wiki page(s).
## Answer
Full answer lives at **[[<answer_doc path without extension>|<doc title>]]**.
It covers:
- <one bullet per major section of the knowledge doc — 3–6 bullets max, each ≤ 12 words>
## Why this matters
<1 sentence — what the user can do with this answer>
> Synthesis: <one line — what kinds of follow-up sources or questions would extend the knowledge doc>Hard rules for the question page:
If the user explicitly flagged a follow-up they want investigated next, ALSO append it to wiki/open-questions.md with the date and cite the question page that spawned it.
After writing both files (knowledge doc + slim question page), regenerate index.md (the new pages change total_wiki_pages):
uv run --script ${CLAUDE_PLUGIN_ROOT:-.claude}/skills/research/scripts/build_index_md.py --research-dir "<research_dir>"
## [YYYY-MM-DD] query | <topic>
- question: "<verbatim, truncated to 200 chars>"
- sources cited: <count>
- wiki pages cited: <count>
- saved as: wiki/questions/<filename> (slim pointer) + <answer_doc path> (knowledge doc — new or enriched) — or "not saved"If the answer wasn't saved, still log it — the log records the conversation flow even when nothing landed in wiki/.
Standard answer formatting:
## <Question rephrased as topic line>
**From research on <topic>** — <N sources cited, M wiki pages>
<answer body with [[wikilinks]] to source pages, entity/concept pages, and where relevant `[Original](<source_url>)` links>
### Sources cited
1. <Title> (origin: <origin>, score: 0.XX) — [[wiki/sources/<slug>]] · [Original](<source_url or "n/a">)
2. ...
<if saved> 📌 Saved: knowledge at `<answer_doc path>` (new / enriched), slim pointer at `wiki/questions/<filename>`. Add to open-questions if you want me to follow up next ingest.
<if not saved> _(Not saved — single-source answer. Tell me "save this" if you want it kept.)_When the caller is another skill (programmatic, not the user directly), drop the conversational framing and return structured data: a YAML-shaped block with sources_cited, wiki_pages_cited, answer_layers_used, saved_question_path, saved_answer_doc_path.
This skill orchestrates external CLIs that may not be installed. Skip this entirely for
query mode. For ingest modes, only check the CLIs that the selected route can actually
use. Seed-only modes (append, seed-only init) check
seed-specific CLIs only: git for GitHub seeds, and no
Obsidian/Readwise/NLM discovery check. Generic web seeds need no CLI — they are fetched
with curl (preinstalled). Discovery modes (deep, init with
discovery) run the full source preflight. A missing CLI must never crash the run with a
cryptic command not found; it must degrade gracefully with a clear, named warning.
For discovery modes, run one detection pass and remember the result as available_clis
plus source_commands (the command string to run for each source CLI). Prefer binaries
on PATH. For Obsidian, also support the installed desktop CLI when it is not on PATH:
obsidian is on PATH, set source_commands.obsidian = "obsidian".OBSIDIAN_CLI is set, use that path.%LOCALAPPDATA%\Programs\Obsidian\Obsidian.com.help; if it says the CLI is not enabled, mark
Obsidian missing and tell the user to enable it in Obsidian Settings > General >
Advanced > Command line interface.For all other CLIs, PATH detection is enough:
for cli in readwise nlm git; do
if command -v "$cli" >/dev/null 2>&1; then
echo "$cli: available"
else
echo "$cli: MISSING"
fi
doneFor seed-only modes, build available_clis from the seed list: include git only if a
GitHub seed is present, and mark other source CLIs as not_needed. Generic web seeds are
fetched with curl and need no preflight.
Degradation policy — apply per source. For every MISSING CLI that is relevant to the selected route, emit a loud one-line warning to the user up front (not silently), naming the CLI, the capability lost, the bundled usage skill, and that the run continues without it:
| CLI | Powers | If MISSING → |
|---|---|---|
obsidian | Obsidian vault search (a research source) | Warn: "⚠️ Obsidian CLI unavailable — skipping your vault as a source. Enable the Obsidian CLI, put obsidian on PATH, or set OBSIDIAN_CLI to the CLI executable. See the obsidian-cli skill." Drop Obsidian from available_clis; continue with other sources. |
readwise | Readwise library + feed search | Warn: "⚠️ readwise CLI not found — skipping Readwise. See the readwise-cli skill (npm install -g @readwise/cli)." Continue. |
nlm | NotebookLM search | Warn: "⚠️ nlm CLI not found — skipping NotebookLM. See the nlm-skill." Set notebook_ids = []; skip Step 3b's auth check. Continue. |
git | GitHub repo ingestion (Step 1a) | Warn (only if the brain dump contains a GitHub URL): "⚠️ git not found — skipping GitHub repo(s): <list>." Skip Step 1a entirely; drop those seeds. Continue. |
YouTube ingestion does not require an API key. It uses public captions via
youtube-transcript-api; videos with unavailable captions are skipped later by the
builder with a clear per-video error.
Hard stop condition. For discovery modes only: if obsidian, readwise, AND nlm
are all MISSING and the brain dump carries no usable seeds (no web URL, no YouTube
URL, no git-able GitHub repo, no local/web PDF), there is nothing to research. Stop and
tell the user clearly:
❌ No research sources available. Install at least one source CLI (
obsidian,readwise, ornlm— see their bundled skills) or include a seed URL/PDF, then re-run.
Otherwise proceed with whatever is available. Pass available_clis and
source_commands into every researcher subagent (Step 4) so it only attempts searches
for installed CLIs and uses the resolved command path when a binary is not on PATH. Auth
(not just presence) is re-checked at point of use for nlm (Step 3b).
Read whatever the user provides. Extract:
If the brain dump contains URIs, process them before moving to step 2:
Vault paths (e.g., Notes/Some Note.md or [[Some Note]]): Read the file directly. These go straight into the research output.
YouTube URLs (youtube.com/watch, youtu.be, youtube.com/shorts, youtube.com/embed, youtube.com/live): Treat as first-class video seeds, not generic web pages. Process them before generic web URLs. Create a seed entry with origin: "youtube", original_path: "youtube://<video_id>", source_url: <url>, youtube_url: <url>, youtube_video_id: <video_id>, transcript_source: "transcript_api", timestamps_available: true, relevance_score: 1.0, and a short placeholder summary from the user's framing. The raw extraction happens in the Builder via scripts/youtube_extract_transcript.py. If captions are unavailable, the Builder records the source as skipped with the script's JSON error unless the user provided a manual transcript file.
Generic web URLs (e.g., https://example.com/article) — any http(s):// link that is not one of the recognized special origins below (not a vault path, not a Readwise reference, not a GitHub repo, not a NotebookLM URI, not a YouTube URL, and not a .pdf): fetch it with curl and strip the HTML to readable text using only preinstalled tools (curl + python3 stdlib — no extra dependency). This handles normal, server-rendered HTML sites; it does not clear bot walls, CAPTCHAs, JS-only rendering, or paywalls — WebFetch is the fallback for those (see below).
curl -fsSL --compressed --max-time 30 \
-A "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120 Safari/537.36" \
"<url>" \
| python3 -c 'import sys,re,html;t=sys.stdin.read();t=re.sub(r"(?is)<(script|style|noscript|template)\b.*?</\1>"," ",t);t=re.sub(r"(?i)</(p|div|h[1-6]|li|tr|section|article|header|footer)>|<br\s*/?>","\n",t);t=re.sub(r"(?s)<[^>]+>"," ",t);t=html.unescape(t);t=re.sub(r"[ \t]+"," ",t);t=re.sub(r"\n[ \t]+","\n",t);t=re.sub(r"\n{3,}","\n\n",t);sys.stdout.write(t.strip())' \
> "working-dir/research-scrape-<slug>.md"curl exits non-zero or the stripped output is suspiciously small (< ~500 chars — typically a JS-only shell or a bot wall), use WebFetch so the run still completes, and tell the user once, clearly: "⚠️ curl couldn't render <url> (likely JS-rendered or bot-walled) — used the lower-fidelity WebFetch fallback." Never let a failed fetch abort the seed.fetched_markdown (see below) so the builder writes the raw file without re-fetching. Set origin: "web", original_path: <full url>, source_url: <full url>.Readwise references: If the user references a specific Readwise source by name, search for it via the MCP tool.
NotebookLM references (e.g., nlm://notebook/<id>, a notebook name, or "my NotebookLM notebook on X"): These reference an entire notebook. Do not treat the notebook itself as a seed source — instead, note the notebook ID for the NLM discovery step (Step 3b) so it's always included in the search. Individual sources within it will be discovered during the research rounds.
GitHub repositories (e.g., https://github.com/owner/repo or .../tree/<branch>): Process via the GitHub pipeline below (Step 1a). Produces one ARCHITECTURE.md always and one <module>.md per targeted module. GitHub never participates in research rounds — it is always seed-only.
Local PDFs (e.g., a path ending in .pdf from the user's filesystem, including paths inside the vault like Media/some-paper.pdf): Treat as a first-class seed. The actual extraction happens in Step 6.2 via scripts/extract_pdf.py — at this stage just record the seed entry with origin: "pdf", original_path: pdf://<basename>, source_url: null, and a placeholder summary (the user's framing if they provided one, else empty — the source_writer will fill it in from extracted text in Step 6.3). Store the absolute path on the seed entry as local_pdf_path so Step 6.2 knows where to extract from.
Web PDFs (a URL ending in .pdf): Same treatment as local PDFs, but Step 6.2 will download to raw/assets/<slug>/original.pdf first via httpx, then extract. If the httpx download is blocked (403 / bot wall / CAPTCHA), retry the download with curl (curl -fsSL --max-time 60 -A "Mozilla/5.0 ... Chrome/120 Safari/537.36" "<url>" -o "raw/assets/<slug>/original.pdf") before giving up.
For each seed URI, create a finding entry (same format as research subagent findings — including author, published_date, publication, source_url metadata fields). For web seed URIs, also include a fetched_markdown field carrying the cleaned content from the curl fetch (or the WebFetch fallback), so the builder can write the file directly without a re-fetch. Seed URIs always get relevance_score: 1.0 — the user explicitly provided them, so they are the highest-relevance sources by definition. They bypass the discovery scoring entirely (they flow through seeds.json, not the discovery results) and are always included in the final output. Never assign a seed URI a score lower than 1.0.
When the brain dump contains a GitHub repo URL, process it BEFORE writing seeds.json.
Guard: requires git. This pipeline shells out to git (via github_clone.py). If
git is MISSING (from Step 0.5), skip the GitHub pipeline entirely — do not attempt
the clone. Warn the user clearly: "⚠️ git not found — skipping GitHub repo(s): <list of repo URLs>. Install git and re-run to include them." Drop those GitHub seeds and continue
with the rest of the ingest.
The GitHub pipeline never clones into the research directory; clones go to a reusable .github-cache/ placed as a sibling of the research dir (i.e. in the research dir's parent — e.g. for Projects/My Project/research-<slug>/ the cache lands at Projects/My Project/.github-cache/). This keeps each project's reusable clones next to its own folder. Only curated spec docs land in the final research dir.
One index entry per repo. Each repo produces a <repo>/ARCHITECTURE.md that acts as a wiki hub plus a set of <repo>/<module>.md neighbor docs. Only ARCHITECTURE.md is registered in index.yaml (uri_full: "<repo>/ARCHITECTURE.md"). The module docs are written to disk alongside it, but they are reached by following links inside ARCHITECTURE — they are not separate entries.
For every unique https://github.com/<owner>/<repo>[...] URL in the brain dump:
Parse targets — run the parser over the brain dump AND any markdown files the user referenced (e.g., an outline.md). The parser extracts repo-relative file paths + line ranges and groups them by parent directory (one "module" per dir):
uv run --script ${CLAUDE_PLUGIN_ROOT:-.claude}/skills/research/scripts/github_parse_targets.py \
--repo "<repo_url>" \
--text "<brain_dump_text>" \
--file "<linked_markdown_file_1>" \
--file "<linked_markdown_file_2>" \
--output "<research_dir>/github-targets-<repo>.json"Pass --text for each inline text blob and --file for each referenced markdown file. If no markdown files were referenced, pass only --text. The output JSON has repo_url, owner, repo, branch, and modules: [...]. An empty modules list means global mode (only ARCHITECTURE.md will be generated).
Shallow-clone the repo into the cache. Pass --research-dir so the cache lands as a sibling of the research dir (its parent), not under working memory:
uv run --script ${CLAUDE_PLUGIN_ROOT:-.claude}/skills/research/scripts/github_clone.py \
--repo "<repo_url>" \
--research-dir "<research_dir>"The script prints {owner, repo, branch, clone_path, commit_sha, action} as JSON on stdout — capture these for the spec writers.
If modules is non-empty, first spawn one github_spec_writer per module IN PARALLEL (module mode) so the architecture writer in the next step can reference real, written files. Each gets:
clone_path, repo_url, owner, repo, commit_sha, branch, research_topicmode: "module"module_path, module_name, files (from the parser output)output_path: <research_dir>/github-staging/<repo>/<module_name>.mdThen spawn ONE github_spec_writer in architecture mode to write <research_dir>/github-staging/<repo>/ARCHITECTURE.md. Pass the same shared inputs plus:
mode: "architecture"output_path: <research_dir>/github-staging/<repo>/ARCHITECTURE.mdmodule_docs: the module entries from Step 1, each {module_path, module_name, filename: "<module_name>.md"}. Empty in global mode.The architecture writer is responsible for:
module_docs entry: a ≤15-word summary + a relative link to the file (e.g., [vectordb](./vectordb.md)).Emit exactly ONE seed entry per repo, keyed to the ARCHITECTURE doc:
origin: "github", relevance_score: 1.0title: "<repo>" (plain, not "<repo> — Architecture" — the ARCHITECTURE is the canonical view of the repo)original_path: "github://<owner>/<repo>@<commit_sha>"source_url: https://github.com/<owner>/<repo>/tree/<commit_sha>github_repo_url: "https://github.com/<owner>/<repo>"github_commit_sha: <commit_sha>github_branch: <branch>github_files: union of every file path referenced across all modules (from the parser output) — used by /research-distill for matching, not for indexing individual docs. Empty list in global mode.authors: ["<owner>"] (augment from README if a clear author is declared)publication: "GitHub"staged_spec_path: absolute path to the staged ARCHITECTURE.md (builder copies from here)Do NOT emit separate entries for module docs — they live alongside ARCHITECTURE in the repo subfolder and are reached via its links.
Once the seed list is assembled (including the GitHub entries), write it to <research_dir>/seeds.json (shape: {"seeds": [ ... ]}) so the Builder Subagent (Step 6) can consume it without going back through the orchestrator's context.
Use the content from seed URIs to enrich your understanding of the topic. Extract additional key themes, terminology, and concepts from them. In discovery modes, these inform the search queries you generate in Step 3; in seed-only modes, they inform the per-source wiki pages and overview/synthesis updates.
Summarize your understanding back to the user in 2-3 sentences so they can correct you if needed.
Skip this step entirely for query, append, and seed-only init. Those modes do not run
research rounds; set rounds_completed: 0 for seed-only ingest.
This step applies to deep mode (and init with explicit discovery). Discovery runs at one
of exactly three depth presets — no free-form round/query counts:
| Preset | Rounds | Queries per round | Total queries |
|---|---|---|---|
fast | 1 | [3] | 3 |
light | 2 | [3, 2] (3 then 2) | 5 |
deep | 3 | [3, 3] (3 each round) | 9 |
Choosing the preset:
fast;
"exhaustive" / "comprehensive" / "thorough" / "deep dive" → deep.fast / light / deep
(default light) before starting rounds.Also capture the topic slug — a short kebab-case name for the output directory (suggest one based on the topic).
Capture the choice as total_rounds and queries_per_round = (round1_count, subsequent_count) per the table above (fast → total_rounds=1, (3,); light →
total_rounds=2, (3, 2); deep → total_rounds=3, (3, 3)), plus topic_slug.
Run this step only for deep or init with explicit discovery. Seed-only modes
skip directly to Step 6 after Step 1.
Based on the brain dump, generate exactly queries_per_round[0] queries (3 for every preset) that approach the topic from different angles. The goal is focused breadth - enough to discover missing context without turning a simple request into a long research run.
Think about:
Write these queries down before spawning subagents — you'll refine them in later rounds.
Run this step only for deep or init with explicit discovery. Seed-only modes do
not query NotebookLM.
The source coverage rule for discovery modes is non-negotiable: every discovery run queries every NotebookLM notebook, every time. No heuristic filtering. This is intentional - we cast the widest possible net only when the user chose deep discovery, and dedup consolidates the findings at the end (nothing is filtered out).
Check presence, then authentication. If nlm was MISSING in Step 0.5, skip this
entire step: warn "⚠️ nlm CLI not found — skipping NotebookLM (see the nlm-skill)",
set notebook_ids = [], and continue. If nlm is present, run nlm login --check. If
that fails, log a loud warning that NotebookLM will be skipped for this run (auth
expired/absent), set notebook_ids = [], and continue. Do NOT silently skip in either
case — the user needs to know NLM was offline so they can decide whether to re-run after
installing/fixing it.
List notebooks: Run nlm notebook list --json to get all available notebooks with their IDs, titles, and source counts.
Filter out empty notebooks only (source_count = 0) — there's nothing to search. Every other notebook is included.
Build notebook_ids: A list of {id, title} objects covering every non-empty notebook. Pass it to every researcher subagent.
User-specified notebooks: If the user referenced specific notebooks in their brain dump (Step 1), they are already in notebook_ids (everything is). Note them in the report so the user knows they were prioritized in the search angles.
Cost trade-off: querying all notebooks scales with the size of the user's NLM library. For libraries above ~30 notebooks this can be slow — accept the cost; the researchers' relevance tags rank the results and dedup consolidates them. If it becomes consistently painful, this is where a --scope knob would land (out of scope for now).
Run this step only for deep or init with explicit discovery.
For each round, spawn one Research Subagent per query in parallel using the Agent tool. Each subagent follows the instructions in agents/researcher.md (read that file and pass its content as the subagent prompt, along with the specific query and context).
The subagent prompt should include:
notebook_ids list from Step 3b (may be empty if NLM is unavailable)available_clis set from Step 0.5 — the researcher searches only the sources whose CLI is available and skips the rest (it must not call a CLI that wasn't listed as available)Subagent result format (saved as JSON by each subagent — see agents/researcher.md Step 5 for the full schema including the metadata fields author, published_date, publication, source_url):
{
"query": "the search query",
"findings": [
{
"title": "Note or highlight title",
"original_path": "full path to source file or readwise identifier",
"origin": "obsidian" or "readwise" or "notebooklm",
"relevance": "high" or "medium",
"summary": "2-3 sentence summary of why this is relevant",
"key_concepts": ["concept1", "concept2"],
"content_preview": "first 200 chars of the content",
"author": "...", "published_date": "...", "publication": "...", "source_url": "..."
}
]
}The researcher caps its output at 15 findings per query — this keeps the candidates JSON small without discarding real coverage.
After each round completes — all steps delegate to bash/scripts/subagents so the round's JSON never enters your context:
Deduplicate the round's subagent outputs via script:
uv run --script ${CLAUDE_PLUGIN_ROOT:-.claude}/skills/research/scripts/dedup_findings.py \
--inputs <research_dir>/round-N-query-*.json \
--output <research_dir>/round-N-deduped.jsonThe script dedups by original_path and keeps the finding with the longest summary. It prints a one-line count — that's all you see.
Generate next-round queries (only if another round follows) by spawning a Gap Analyzer Subagent. Read agents/gap_analyzer.md and pass its content as the subagent prompt, along with:
research_topic, input_summary, key_themes from Step 1deduped_findings_path: <research_dir>/round-N-deduped.jsonround_number: N (the round that just completed)total_rounds: the configured totalprevious_queries: queries from all prior rounds (you have these — they're what you generated/dispatched)target_query_count: queries_per_round[1] (default 3)output_path: <research_dir>/round-{N+1}-queries.jsonThe subagent returns only the path to its output JSON. You then read just the queries[].query field (via jq -r '.queries[].query') to pick up the next round's queries. Never load the full analyzer output into your context.
Continue for N rounds (as configured in step 2).
Skip this step for append and seed-only init. For those routes, write
<research_dir>/discovery-results.json as {"results": []} and continue to Step 6 with
seeds.json only.
There is no LLM reranker. The research rounds already cast a wide net and tagged each finding
relevance: high|medium; this step just consolidates them deterministically before any wiki
element is computed. It is a single non-LLM script call — nothing here enters your context.
Merge all per-round deduped files into one scored results file. The same dedup script
that ran per-round handles cross-round deduplication; the --as-results flag wraps the
output as {"results": [...]} and derives a numeric relevance_score from each finding's
relevance tag (high → 0.8, medium → 0.5, else 0.5):
uv run --script ${CLAUDE_PLUGIN_ROOT:-.claude}/skills/research/scripts/dedup_findings.py \
--inputs <research_dir>/round-*-deduped.json \
--as-results \
--output <research_dir>/discovery-results.jsonEvery candidate is kept (no filtering); deduplication is by original_path, keeping the
finding with the longest summary. The researcher's broad summaries carry through to
index.yaml; the per-source wiki page (Layer 1.5, written in Step 6.3) is where the
sharper, intent-aligned summary lives.
In parallel with this merge, kick off the seed-URI portion of the Builder Subagent
(Step 6). Seed URIs (score 1.0) need no scoring, so their files can be copied/fetched
immediately. Pass the builder a seeds-only variant with results_json = empty results and
seeds_json = your seed list. Join with the full builder in Step 6.
Use discovery-results.json for the full Step 6 builder run. Sorting is handled
deterministically by build_index_yaml.py (seeds first, then relevance_score
descending) — you do not sort by hand.
The relevance_score in index.yaml tells future agents which sources to prioritize, but
nothing is discarded.
The Builder Subagent copies / fetches / pipes raw source content onto disk under <research_dir>/raw/. It does NOT build index.yaml — that happens in Step 6.5 after the wiki layer is in place, so the index can include uri_source_page references. You, the orchestrator, should never see individual source file contents in Step 6.
Compute the research-dir path. Default: working-dir/research-<topic-slug>/. Override if the user specified a different output path. Create the v4 layout up front:
mkdir -p "<research_dir>/raw/assets" "<research_dir>/wiki"Spawn the Builder Subagent. Read agents/builder.md and pass it as the subagent prompt, along with:
results_json: <research_dir>/discovery-results.jsonseeds_json: <research_dir>/seeds.json (from Step 1)research_dir: the path from step 1 above (final files go here AND scratch JSONs land here too — they're cleaned up in Step 7)topic, input_summary, rounds_completed from Step 1 / Step 2skill_dir: ${CLAUDE_PLUGIN_ROOT:-.claude}/skills/research (absolute path)mode: "raw_only" — explicitly tells the builder to skip index.yaml emissionIn discovery modes, if you already started the seed-only builder in Step 5, wait for it to finish before launching the full builder - the full run is idempotent. In seed-only modes, launch the builder once with empty discovery results and the complete seeds.json.
The builder returns:
{"built": 25, "skipped": 2, "research_dir": "/path/to/research-<slug>", "raw_files": [...], "skipped_details": [...]}raw_files is a list of {source_idx, raw_path, original_path, slug, origin} entries — one per successful source. You'll feed this to Step 6.2 (asset processing) and Step 6.3 (source-page writing). skipped_details (truncated to 10) lists sources that fetched no content.
Layer 2 (uri_highlights) rule unchanged: populated only when the source carries manually user-curated highlights (Readwise highlights synced into the vault at Sources/Readwise/, and raw Readwise documents whose vault copy holds the user's highlights). For everything else, uri_highlights is null and the full content lives at uri_full (Layer 3).
For each entry in raw_files, run the asset pipeline. All bash; nothing here goes through your LLM context.
PDFs — for any source where origin == "pdf" (user-dropped) or where the original is a .pdf URL, the Builder will have already invoked extract_pdf.py to extract markdown + images and preserve the original at raw/assets/<slug>/original.pdf. If somehow not done, run it now:
uv run --script ${CLAUDE_PLUGIN_ROOT:-.claude}/skills/research/scripts/extract_pdf.py \
--pdf "<original_pdf_path>" \
--output-md "<research_dir>/raw/<slug>.md" \
--assets-dir "<research_dir>/raw/assets/<slug>"Images — for every successful raw markdown file, download referenced images to local assets and rewrite refs:
uv run --script ${CLAUDE_PLUGIN_ROOT:-.claude}/skills/research/scripts/download_assets.py \
--markdown "<research_dir>/raw/<slug>.md" \
--assets-dir "<research_dir>/raw/assets/<slug>"The script is a noop when there are no remote image references. Capture the JSON it prints; the downloaded array becomes the source's assets field in the index.
Build a per-source assets list — the manifest of asset files actually present on disk under raw/assets/<slug>/. This goes into the index in Step 6.5.
You can run the asset pipeline for many sources in parallel via a bash loop or xargs -P. Failures (e.g. a 404 on an image) are non-fatal — they're logged in the JSON output and the script leaves the original URL in place.
For each entry in raw_files, spawn one source_writer subagent (see agents/source_writer.md) in parallel using a single message with multiple Agent calls. Each subagent writes exactly one wiki/sources/<slug>.md. You never read the raw files yourself — that's the source_writer's job.
Each spawn passes:
raw_path: absolute path to the raw filesource_metadata: the JSON object for this source from the discovery results (already has all the index fields, including relevance_score)research_topic, input_summary (from Step 1)existing_entities: list of {slug, name} for entity pages already in wiki/entities/. On init runs this is empty; on append runs it's populated by ls-and-frontmatter-extract via bash:for f in <research_dir>/wiki/entities/*.md; do
awk '/^---$/{c++; next} c==1 && /^name:/' "$f"
doneexisting_concepts: same pattern for wiki/concepts/assets: the per-source asset list from Step 6.2output_path: <research_dir>/wiki/sources/<slug>.mdCollect the JSON outputs from each subagent. Aggregate them into a single in-memory structure (or write to <research_dir>/source-writer-outputs.json). You'll need:
entities_referenced and concepts_referenced — counts per slug across sources (used in Step 6.4 to decide which entity/concept pages to write).suggested_new_entities and suggested_new_concepts — candidates for new pages.Before spawning the entity/concept writers in Step 6.4, surface the emerging shape to the user and let them redirect emphasis. Skip this step entirely for append and seed-only init unless the user explicitly asks for a discussion checkpoint. Skip it also when the user says "skip discussion" up front or when ingest passes touched <= 3 new sources - the cost outweighs the benefit at small scale.
Compute and present (to the user, not via subagent):
Use AskUserQuestion with the four buckets above as a single multi-select question:
Before writing entity/concept pages, anything to redirect?
- ☐ Drop source: <title> (it's tangential to the topic)
- ☐ Merge slugs: "ralph-loop" + "ralph_loop" → "ralph-loop"
- ☐ Add comparison: <a> vs <b>
- ☐ Save these open questions: <list>
- ☐ Proceed as-isAlso offer "Other" so the user can type free-form redirects ("emphasize agent-loop more in synthesis", "demote source X to score 0.3", etc.). Apply the user's edits to your in-memory aggregate before spawning Step 6.4 writers. Source pages that have already been written stay; you just don't propagate their entities/concepts to the wiki layer.
Discussion is opt-out — the user can always say "skip discussion" to bypass this step. But default it ON. The whole point of Karpathy's pattern is the conversational loop; this step is where it lives.
Aggregate slug → mention counts across all source_writer outputs. Apply the threshold: an entity or concept page is created (or updated) when mention_count >= 2 across distinct sources. Single-source mentions stay only on the source page.
For each qualifying slug, spawn one wiki_page_writer subagent (see agents/wiki_page_writer.md) in parallel. Pass:
page_type: "entity" or "concept"slug, name, aliases (from the source_writer outputs, deduped)source_pages: list of absolute paths to the wiki/sources/*.md pages that mention this slug. Only these — the writer reads only its in-scope source pages.research_topic, input_summaryexisting_page_path: path to existing wiki/entities/<slug>.md or wiki/concepts/<slug>.md if one exists; null otherwiseoutput_path: where to writeCollect the JSON outputs. Each tells you {action: created|updated|noop, source_count, confidence, tensions_found, open_questions_surfaced}. Aggregate tensions_found > 0 as a flag for the lint pass; aggregate open_questions_surfaced for Step 6.5.
You can run dozens of these in parallel — each is bounded by its source_pages list and reads no raw files.
Spawn two wiki_summary_writer subagents (see agents/wiki_summary_writer.md) in parallel:
page_type: "overview", output_path: <research_dir>/wiki/overview.mdpage_type: "synthesis", output_path: <research_dir>/wiki/synthesis.mdBoth receive research_dir, research_topic, input_summary, and existing_page_path (if a prior overview/synthesis exists). They discover wiki state themselves via bash and respect a context budget — you don't pre-load wiki pages for them.
Aggregate any open_questions_surfaced reported by Step 6.4 wiki_page_writer outputs. Append them to <research_dir>/wiki/open-questions.md with a date-stamped section:
cat >> "<research_dir>/wiki/open-questions.md" <<EOF
## [$(date -u +%Y-%m-%d)] from ingest of $(wc -l < <(echo "<raw_files>")) source(s)
- <question 1>
- <question 2>
EOFIf open-questions.md doesn't yet exist, create it with a header line: # Open questions\n. Lint will surface and prune these later.
Now that the wiki layer exists, build the canonical index. Two scripts, in order:
Augment discovery results and seeds with uri_source_page and assets for each source (via jq — never load JSON in your context):
# For each source in discovery-results.json, add uri_source_page and assets fields
# using the per-source mapping you built in Step 6.3.
jq --slurpfile aug "<research_dir>/source-augments.json" \
'.results |= map(. + ($aug[0][.original_path] // {}))' \
"<research_dir>/discovery-results.json" \
> "<research_dir>/discovery-final.json"
# Do the same for seed sources; seed-only modes rely on this file.
jq --slurpfile aug "<research_dir>/source-augments.json" \
'.seeds |= map(. + ($aug[0][.original_path] // {}))' \
"<research_dir>/seeds.json" \
> "<research_dir>/seeds-augmented.json"Where <research_dir>/source-augments.json is a {<original_path>: {uri_source_page: "...", assets: [...]}} map you assembled from source_writer + asset-pipeline outputs.
Run build_index_yaml.py to emit the canonical index:
PRIOR_CREATED=$(grep '^created:' "<research_dir>/index.yaml" 2>/dev/null | awk -F"'" '{print $2}' || true)
# APPEND ONLY: snapshot the prior index so its existing sources are merged forward.
# Skip this on init (no prior index). The snapshot avoids any read/write aliasing on
# the same path and is removed in Step 7's *.json/scratch cleanup pass.
PRIOR_INDEX=""
if [ -f "<research_dir>/index.yaml" ]; then
PRIOR_INDEX="<research_dir>/.prior-index.yaml"
cp "<research_dir>/index.yaml" "$PRIOR_INDEX"
fi
uv run --script ${CLAUDE_PLUGIN_ROOT:-.claude}/skills/research/scripts/build_index_yaml.py \
--results "<research_dir>/discovery-final.json" \
--seeds "<research_dir>/seeds-augmented.json" \
--research-dir "<research_dir>" \
--topic "<topic>" \
--input-summary "<input_summary>" \
--rounds <rounds_completed> \
${PRIOR_CREATED:+--existing-created "$PRIOR_CREATED"} \
${PRIOR_INDEX:+--existing-index "$PRIOR_INDEX"} \
--output "<research_dir>/index.yaml"
rm -f "<research_dir>/.prior-index.yaml"On init runs there is no prior index.yaml, so PRIOR_CREATED/PRIOR_INDEX are
empty and the index is built fresh from this run's seeds + results. On append and
deep runs against an existing dir, the prior index is snapshotted and
passed via --existing-index, so its existing sources are carried forward and unioned
with the new ones (deduped by original_path, new wins) — without this, an append
would rebuild the index from only the newly-ingested sources and drop every
pre-existing source. --existing-created preserves the original created timestamp.
Run build_index_md.py to regenerate the Obsidian-readable view:
uv run --script ${CLAUDE_PLUGIN_ROOT:-.claude}/skills/research/scripts/build_index_md.py \
--research-dir "<research_dir>"This is byte-stable: identical inputs → byte-identical output. Always the last write before the log entry.
Append a structured entry. Format is non-negotiable (the ## [YYYY-MM-DD] <op> | <subject> prefix is what makes the log greppable):
DATE=$(date -u +%Y-%m-%d)
cat >> "<research_dir>/log.md" <<EOF
## [$DATE] ingest | <topic> ($(echo <raw_files_count>) sources)
- mode: <append | deep | init>
- rounds: <rounds_completed>
- expected runtime shown: <rough estimate from routing plan>
- raw files added: <count>
- wiki pages written: <count> sources, <count> entities, <count> concepts, overview + synthesis updated
- open questions surfaced: <count>
- skipped: <count> ($(echo <skipped_titles | head -3>))
EOFIf log.md doesn't exist, create it with a # Log\n header first.
After the research directory is built successfully, delete intermediary JSONs that live alongside the canonical artifacts (top-level only — raw/ and wiki/ are recursed into for nothing here):
round*-query*.jsonround*-deduped.jsonall-rounds-deduped.jsondiscovery-results.json, discovery-final.json, and discovery-with-uris.jsonseeds.json, seeds-augmented.json, seeds-with-uris.jsonsource-augments.jsonsource-writer-outputs.jsonnotebook-ids.jsonrm -f "<research_dir>"/*.jsonThe glob is safe: index.yaml is YAML (not JSON), and raw//wiki/ are subdirs the glob does not descend into. After cleanup, only index.yaml, index.md, log.md, raw/, and wiki/ remain at the research-dir root.
Tell the user:
query, append, deep, or init) and whether discovery ran (and at which depth preset)wiki/overview.md and wiki/synthesis.md"wiki/open-questions.mdskipped > 0, one-line summaryindex.md in Obsidian (Obsidian-readable navigation)wiki/synthesis.md (the thesis) and wiki/overview.md (the catalog)index.yaml to /research (query mode), which handles progressive disclosure/research-lint to health-check the wiki, or /research-render marp <topic> to export a slide deck."Search strategy — Obsidian: Use the obsidian CLI for search operations: obsidian search query="<terms>" limit=20. Once you have the file paths from search results, read the files directly using the Read tool. See the obsidian-cli skill for full command reference. Guard: only run this if obsidian is in available_clis (presence checked in Step 0.5 via command -v obsidian); if it's missing, skip the vault as a source — never let command not found abort the search.
Search strategy — Readwise: Use the readwise CLI (not the MCP tool). For the full command reference, access the readwise-cli skill. Guard: only run this if readwise is in available_clis (presence checked in Step 0.5 via command -v readwise); if it's missing, skip Readwise as a source with the Step 0.5 warning — never abort on command not found. Readwise splits into two areas and research must cover both:
reader-search-documents (new, later, shortlist, archive).--location-in feed.Key commands the researcher subagent uses:
readwise reader-search-documents --query "<query>" — library document search (default).readwise reader-search-documents --query "<query>" --location-in feed — feed document search.readwise reader-search-documents --query "<query>" --note-search "<query>" — searches the user's own annotations written on documents (high signal). Run a feed-scoped variant too.readwise readwise-search-highlights --vector-search-term "<query>" — semantic search across all highlights (spans both areas).readwise reader-get-document-details --document-id <id> — full document content as Markdown, used for the two-layer pattern.The researcher subagent tags each Readwise finding with readwise_location: "library" | "feed" so library hits can be weighted slightly higher than feed hits when two sources look otherwise equivalent, and so future agents can filter by location.
File naming: When copying files to the research dir, slugify titles to kebab-case, remove special characters, and keep filenames under 60 characters. For Readwise highlights, prefix with readwise-.
Deduplication matters: The same note can be found by multiple queries. Always deduplicate by original_path before building the output. For NotebookLM sources that have a source_url, also deduplicate against Readwise and web sources sharing the same URL — prefer the Readwise version (it carries user-curated highlights and annotations).
Search strategy — NotebookLM: Use the nlm CLI. Key commands for research:
nlm login --check — verify authentication before searchingnlm notebook list --json — list all notebooks with IDs and titlesnlm note list <nb-id> — list user-created notes in a notebook (these are prioritized over raw sources)nlm notebook query <id> "question" — one-shot Q&A against a notebook's sources (returns AI answer with cited sources)nlm source list <nb-id> — list imported sources in a notebook with IDs and metadatanlm source describe <source-id> — AI summary and keywords for a sourcenlm source content <source-id> — raw text content of a source
Sessions expire in ~20 minutes. If commands fail with auth errors, skip NLM for the remainder of the research run.NotebookLM note vs source priority: Notebooks contain two types of content — notes (user-written synthesis, right panel in the UI) and sources (imported articles/papers, left panel). Notes always take priority because the user deliberately created them. The researcher subagent fetches notes first via nlm note list, then uses nlm notebook query to discover relevant raw sources. Notes get nlm_content_type: "note" and use nlm://note/<id> paths; raw sources get nlm_content_type: "source" and use nlm://source/<id> paths.
Extensibility: This skill searches Obsidian, Readwise, and NotebookLM, and accepts web URLs + GitHub repos as seed-only URIs (GitHub handled via the Step 1a pipeline). Generic web links are fetched with curl + the stdlib HTML stripper (WebFetch fallback) — see Step 1, point 3. The architecture is designed so new sources can be added by extending the researcher agent with additional search steps; a natural next step is web-search rounds (the researcher harvests result URLs from a search provider, then fetches the promising ones with the same curl recipe). The orchestrator doesn't need to change — it just spawns researchers and collects results.
agents/researcher.md — Research Subagents (one per query per round). Read and include its content when spawning each research subagent.agents/gap_analyzer.md — Gap Analyzer Subagent. Spawned between rounds to generate the next round's queries from the previous round's deduped findings.agents/builder.md — Builder Subagent. Spawned in Step 6 to produce raw files (no longer writes index.yaml — that moved to Step 6.7 after wiki pages exist).agents/github_spec_writer.md — GitHub Spec Writer Subagent. Spawned in Step 1a — once in architecture mode per repo, plus one per targeted module.agents/source_writer.md — Source Writer Subagent. Spawned in Step 6.3, one per raw source in parallel. Reads one raw file and writes wiki/sources/<slug>.md with extended summary, key claims, quotes, and connections; returns suggested entities/concepts.agents/wiki_page_writer.md — Wiki Page Writer Subagent (entity OR concept). Spawned in Step 6.4, one per qualifying slug in parallel. Aggregates from per-source pages only — never reads raw.agents/wiki_summary_writer.md — Wiki Summary Writer Subagent (overview OR synthesis). Spawned in Step 6.5, two in parallel. Discovers wiki state itself via bash; capped at 5 full reads.scripts/dedup_findings.py — Deduplicates research findings by original_path, keeping the entry with the longest summary. Used per-round; with --as-results it also merges all rounds into the scored discovery-results.json (deriving relevance_score from each finding's relevance tag) that the builder and index consume — this replaces the old reranker pass.scripts/build_index_yaml.py — Emits canonical index.yaml from the discovery results plus seed URIs. Handles schema, sort order, and created/last_updated semantics deterministically. Called in Step 6.7 after the wiki layer is built so uri_source_page and assets fields can be populated.scripts/build_index_md.py — Generates Obsidian-readable index.md from index.yaml. Idempotent + byte-stable. Always the last write before the log entry.scripts/extract_pdf.py — PDF → markdown via pymupdf4llm. Extracts images to raw/assets/<slug>/, preserves the original PDF, flags low-quality (scanned) PDFs. Used in Step 6.2 for user-dropped or web-fetched PDFs.scripts/download_assets.py — Scans a markdown file for remote image references, downloads them to raw/assets/<slug>/, and rewrites references to local paths. Used in Step 6.2 on every raw markdown file.scripts/github_parse_targets.py — Parses the brain dump (and referenced markdown files) for GitHub file references against a specific repo. Groups them into modules by parent directory. Used in Step 1a.scripts/github_clone.py — Shallow-clones a GitHub repo into a reusable .github-cache/ placed as a sibling of the research dir (--research-dir → its parent) and returns the HEAD SHA. Used in Step 1a.© iusztinpaul, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 17 other files (scripts) in plugins/ai-research-os/skills/research of iusztinpaul/ai-research-os-workshop.
Open the folder on GitHubat commit dc66605
Research next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Research this skilliusztinpaul/ai-research-os-workshop | 179 | — | ~17k | Automated safety check: Pass | MIT | |
| Multi-Source to NotebookLM Processorjoeseesun/qiaomu-anything-to-notebooklm | 6.2k | — | ~3.6k | Automated safety check: Pass | MIT | |
| NotebookLM Research Workflowclaude-world/notebooklm-skill | 467 | — | ~1.8k | Automated safety check: Pass | MIT | |
| NotebookLM CLI Guidejacob-bd/notebooklm-cli | 256 | — | ~3.4k | Automated safety check: Warn | MIT | |
| Deep Research Notebooklmdavila7/claude-code-templates | 32k | — | ~1.7k | Automated safety check: Pass | MIT | |
| Notebooklmalirezarezvani/claude-skills | 28k | — | ~4k | Automated safety check: Pass | MIT |
joeseesun/qiaomu-anything-to-notebooklm
Collects content from WeChat articles, web pages, YouTube, podcasts, documents and more, uploads it to NotebookLM and generates podcasts, slides or mind maps.
claude-world/notebooklm-skill
Creates NotebookLM notebooks from URLs, text and files, asks cited questions, runs web research and generates audio, slides, quizzes and other artifacts.
jacob-bd/notebooklm-cli
Guides use of the nlm command-line tool to automate Google NotebookLM: notebooks, sources, research, one-shot questions and generated podcasts, reports, quizzes and slides.
davila7/claude-code-templates
Deep research skill powered by NotebookLM MCP. An agent skill from davila7/claude-code-templates.
alirezarezvani/claude-skills
Browser automation skill for controlling Google's NotebookLM.
PleasePrompto/notebooklm-skill
Lets Claude Code ask questions of your Google NotebookLM notebooks through browser automation and return answers grounded in your uploaded sources.
iusztinpaul/ai-research-os-workshop
Health-check a research directory produced by /research. An agent skill from iusztinpaul/ai-research-os-workshop.
iusztinpaul/ai-research-os-workshop
Expert guide for the NotebookLM CLI (nlm) and MCP server - interfaces for Google NotebookLM.
iusztinpaul/ai-research-os-workshop
Distill a research directory (produced by /research) into a single compact research.md containing a guideline-relative distillation of only the sources that were actually used in a piece of content.
iusztinpaul/ai-research-os-workshop
Generate a multi-form answer (Marp slide deck, matplotlib chart, Obsidian Canvas, or social content brief) from one or more wiki pages in a research directory and file the output back into…
iusztinpaul/ai-research-os-workshop
How to use the Readwise CLI — access highlights, documents, and your entire reading library from the command line
Works with
Build, extend, AND query a persistent LLM-maintained wiki for any research topic. Research is an agent skill from iusztinpaul/ai-research-os-workshop. Build, extend, AND query a persistent LLM-maintained wiki for any research topic.
Research fits situations like: any research interaction - first-time research on a topic; what do I have on X; load my research on Y; add this PDF to my research.
Run `npx skills add iusztinpaul/ai-research-os-workshop --skill research -a claude-code`. Or copy the skill folder (plugins/ai-research-os/skills/research in iusztinpaul/ai-research-os-workshop) into .claude/skills/research in your project. Claude Code loads it when a task matches its description.
Run `npx skills add iusztinpaul/ai-research-os-workshop --skill research -a codex`. Or copy the skill folder (plugins/ai-research-os/skills/research in iusztinpaul/ai-research-os-workshop) into .agents/skills/research in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add iusztinpaul/ai-research-os-workshop --skill research -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/research, .gemini/skills/research, .github/skills/research and .opencode/skills/research in your project.
Going by SKILL.md and its folder, Research needs Python for the scripts in its folder and the command-line tools its instructions call (uv, jq, curl, python3 and npm). Our summary lists: Python 3.
SKILL.md names 1 domain. In commands or code: github.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Research is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 17k tokens (SKILL.md is roughly 69k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Research: Multi-Source to NotebookLM Processor (joeseesun/qiaomu-anything-to-notebooklm, 6.2k stars), NotebookLM Research Workflow (claude-world/notebooklm-skill, 467 stars), NotebookLM CLI Guide (jacob-bd/notebooklm-cli, 256 stars) and Deep Research Notebooklm (davila7/claude-code-templates, 32k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
iusztinpaul (a GitHub user) maintains it in iusztinpaul/ai-research-os-workshop, which has 179 GitHub stars. The repository holds 6 skills in this directory. The repository was last updated on June 27, 2026.
Source: iusztinpaul/ai-research-os-workshop on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.