RAG Cite Sources
lyonzin/knowledge-rag
Every technical claim drawn from the local corpus must ship with a source citation formatted as path:line or path:section.
Sciverse academic paper retrieval: structured metadata search, semantic chunk retrieval for RAG, and character-range content reading (offsets in Unicode code points).
$ npx skills add opendatalab/Sciverse-Agent-Tools --skill sciverse-academic-retrieval -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install opendatalab/Sciverse-Agent-Tools sciverse-academic-retrieval --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/opendatalab/Sciverse-Agent-Tools.git skills-src && mkdir -p .claude/skills && cp -r skills-src/clawhub .claude/skills/sciverse-academic-retrieval && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "sciverse-academic-retrieval" agent skill from https://github.com/opendatalab/Sciverse-Agent-Tools/tree/main/clawhub into .claude/skills/sciverse-academic-retrieval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "sciverse-academic-retrieval", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/opendatalab/Sciverse-Agent-Tools/tree/main/clawhubType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add opendatalab/Sciverse-Agent-Tools --skill sciverse-academic-retrieval -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install opendatalab/Sciverse-Agent-Tools sciverse-academic-retrieval --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/opendatalab/Sciverse-Agent-Tools.git skills-src && mkdir -p .agents/skills && cp -r skills-src/clawhub .agents/skills/sciverse-academic-retrieval && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "sciverse-academic-retrieval" agent skill from https://github.com/opendatalab/Sciverse-Agent-Tools/tree/main/clawhub into .agents/skills/sciverse-academic-retrieval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "sciverse-academic-retrieval", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add opendatalab/Sciverse-Agent-Tools --skill sciverse-academic-retrieval -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install opendatalab/Sciverse-Agent-Tools sciverse-academic-retrieval --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/opendatalab/Sciverse-Agent-Tools.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/clawhub .cursor/skills/sciverse-academic-retrieval && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "sciverse-academic-retrieval" agent skill from https://github.com/opendatalab/Sciverse-Agent-Tools/tree/main/clawhub into .cursor/skills/sciverse-academic-retrieval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "sciverse-academic-retrieval", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/opendatalab/Sciverse-Agent-Tools.git --path clawhub--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add opendatalab/Sciverse-Agent-Tools --skill sciverse-academic-retrieval -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install opendatalab/Sciverse-Agent-Tools sciverse-academic-retrieval --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/opendatalab/Sciverse-Agent-Tools.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/clawhub .gemini/skills/sciverse-academic-retrieval && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "sciverse-academic-retrieval" agent skill from https://github.com/opendatalab/Sciverse-Agent-Tools/tree/main/clawhub into .gemini/skills/sciverse-academic-retrieval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "sciverse-academic-retrieval", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install opendatalab/Sciverse-Agent-Tools sciverse-academic-retrievalInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add opendatalab/Sciverse-Agent-Tools --skill sciverse-academic-retrieval -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/opendatalab/Sciverse-Agent-Tools.git skills-src && mkdir -p .github/skills && cp -r skills-src/clawhub .github/skills/sciverse-academic-retrieval && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "sciverse-academic-retrieval" agent skill from https://github.com/opendatalab/Sciverse-Agent-Tools/tree/main/clawhub into .github/skills/sciverse-academic-retrieval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "sciverse-academic-retrieval", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add opendatalab/Sciverse-Agent-Tools --skill sciverse-academic-retrieval -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install opendatalab/Sciverse-Agent-Tools sciverse-academic-retrieval --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/opendatalab/Sciverse-Agent-Tools.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/clawhub .opencode/skills/sciverse-academic-retrieval && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "sciverse-academic-retrieval" agent skill from https://github.com/opendatalab/Sciverse-Agent-Tools/tree/main/clawhub into .opencode/skills/sciverse-academic-retrieval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "sciverse-academic-retrieval", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
sciverse-academic-retrievalSciverse academic paper retrieval: structured metadata search, semantic chunk retrieval for RAG, and character-range content reading (offsets in Unicode code points).
Sciverse Academic Retrieval is an agent skill from opendatalab/Sciverse-Agent-Tools. Sciverse academic paper retrieval: structured metadata search, semantic chunk retrieval for RAG, and character-range content reading (offsets in Unicode code points). For agent workflows that need citation-grade scientific literature.
Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 10 other files, including scripts (for example `README.md` and `manifest.json`).
It sits in AI & LLM Engineering, covering Citation management and Retrieval-augmented generation. The repository describes itself as: Standardized tool schemas and SDKs that expose Sciverse Open Platform retrieval capabilities to LLM agents. The licence is Apache-2.0.
Read from SKILL.md and the folder at commit 5246a81. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 7 files in scripts/ (JavaScript), which the agent can run.
Shell commands in SKILL.md call:
nodeFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
sciverse.spaceFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
SCIVERSE_API_TOKENFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Sciverse Academic Retrieval loads about 2.3k tokens when it runs. Until then it costs about 66 tokens; SKILL.md has 918 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from opendatalab/Sciverse-Agent-Tools at commit 5246a81, republished under its Apache-2.0 licence (© opendatalab). 918 words, ~2,323 tokens.
.claude/skills/sciverse-academic-retrieval/SKILL.md (or your agent's skills folder). This skill also uses 9 other files; get the full folder from GitHub.Sciverse academic paper retrieval: structured metadata search, semantic chunk retrieval for RAG, and character-range content reading (offsets in Unicode code points). For agent workflows that need citation-grade scientific literature.
Trigger this skill when the user's request involves any of:
This skill requires the SCIVERSE_API_TOKEN environment variable
(obtain from https://sciverse.space). Optionally set SCIVERSE_BASE_URL
to override the default API base URL.
Search academic papers by structured filters (title, authors, journal, year, subjects, etc.). Use when: "find Hinton's papers from 2020-2023", "Nature papers on CRISPR". Not for: natural-language Q&A retrieval (use semantic_search) or full-text snippets (use read_content). Returns: list of papers; each entry has unique_id (always present), doc_id (only when full text exists), title, author, abstract, publication_venue_name_unified, publication_published_year.
Invoke: node scripts/search_papers.mjs '<JSON args>'
Natural-language semantic search returning relevant paper chunks for RAG-style answering. Use when: "How does Transformer attention work?", "What are recent methods for protein structure prediction?". Not for: precise field filtering (use search_papers) or fetching full original text (use read_content). Returns: list of chunks; each entry has chunk_id, doc_id, abstract, chunk, score, title, offset. Typical chain: semantic_search → pick chunk → read_content(doc_id, offset).
Invoke: node scripts/semantic_search.mjs '<JSON args>'
Returns the schema catalog for search_papers: every field name, type, whether it's filterable / sortable, default-return status, human description, and applicable FilterOperators. Use when: "Which field do I filter by DOI?", "What values can access_oa_status take?", "What's the right enum for metadata_type?". Not for: actually searching papers (use search_papers / semantic_search). Typical pattern: call once when first encountering Sciverse or facing an ambiguous field need, then construct precise search_papers filters from the returned schema. Pass include_sample_values=true to also fetch top-20 values for enum-like fields (OpenSearch terms aggregation, 24h cached).
Invoke: node scripts/list_catalog.mjs '<JSON args>'
Paginate the full relation list of a paper. citations/references/related_works are unbounded arrays (up to 340k entries for a single paper) and are NOT projectable in search_papers, so this endpoint is the only way to read them. Use when: "What does paper X cite?" (relation=REFERENCES), "Which papers cite paper X?" (relation=CITATIONS), "Works related to paper X" (relation=RELATED_WORKS). Note: CITATIONS (incoming: who cites me) and REFERENCES (outgoing: who I cite) are opposite directions. Typical chain: get unique_id from search_papers / semantic_search, then paginate here by relation. Two limits (CITATIONS only; REFERENCES/RELATED_WORKS max out at 11833/20 in practice): more than 10000 relations returns 429; page*page_size above 10000 returns 400. In both cases switch to search_papers with filters_advanced on references_unique_id — it supports deep paging and arbitrary sorting. total_count counts in-corpus matches only, so it can differ from the paper's own citation_count by about 1%.
Invoke: node scripts/list_paper_relations.mjs '<JSON args>'
Read a range of a paper's original text addressed in Unicode code points (offset/limit count characters like Python len(), not bytes). Typically used with a doc_id/offset returned by semantic_search to expand context (read more text before or after a chunk). Returns: text fragment, bytes_returned (UTF-8 byte length of text, for reference only), next_offset (code-point offset of the next fragment — page with it, never with bytes_returned), more (boolean). Server behaviour: limit above 524288 is silently clamped; omitting offset returns the whole document ignoring limit — the SDKs / MCP server send offset=0 and limit=4096 by default, so pass offset explicitly when calling the HTTP API directly.
Invoke: node scripts/read_content.mjs '<JSON args>'
Returns the binary bytes of a paper figure / table image referenced
inside read_content's Markdown via  placeholders.
Use when the user asks to see / display / describe a figure and
read_content output contains an image reference.
Input file_name comes from the Markdown URL part (relative path,
no \\ or ..).
Returns: raw image stream + image/* Content-Type. The SDK / MCP
server wraps the bytes as base64 + mimeType so Claude (multimodal)
can read the image directly.
Invoke: node scripts/get_resource.mjs '<JSON args>'
If you're unsure which fields exist or what values an enum takes
(e.g. metadata_type, language, access_oa_status), call
list_catalog once at the start. Sample values are returned for
low-cardinality fields. Use it instead of guessing field names —
guessing wastes turns.
list_catalog(include_sample_values=true)
└─▶ fields[].name + sample_values → precise filter constructionRAG flow (natural-language Q&A):
semantic_search(query=...) → hits[i].doc_id, hits[i].offset
└─▶ read_content(doc_id, offset)Lookup by DOI:
search_papers(filters_advanced=[{field: "doi", value: "10.1038/..."}])OA + year filter:
search_papers(
year_from=2024,
filters_advanced=[{field: "access_is_oa", value: "true"}]
)Scoped semantic search (constrained corpus):
semantic_search(
query="...",
filters={"author": ["Hinton"],
"publication_published_year": {"gte": 2020}}
) # applied at recall time, server-side; AND across fieldsSoft semantics: chunks missing that metadata are NOT excluded. For a hard guarantee, or meta-only constraints (fwci, citation graph, complex hit-sets), scope by doc_id — a HARD recall-time filter:
search_papers(..., fields=["doc_id","title"]) → collect doc_id
semantic_search(query=..., filters={"doc_id": [...]})
# hits never leave the set; empty list → empty hits (never global);
# up to 1000 deduped ids (400 SCOPE_TOO_LARGE beyond)Bias fuzzy search ranking (soft boosts — stackable):
Three multiplicative boosts (freshness_boost / impact_boost /
language_affinity, each NONE/MILD/STRONG) reorder fuzzy-search
results while keeping relevance. Only effective when query is
non-empty; ignored when any sort is set; shallow paging while active.
sort_by_year defaults to auto (relevance with query, newest-first
for pure filters); query+desc is an anti-pattern — it degrades the
query to a match filter and disables all boosts; use freshness_boost.
search_papers(query="large language model", freshness_boost="STRONG")
# recent first: STRONG=3-year decay, MILD=10-year
search_papers(query="protein folding", impact_boost="MILD")
# highly-cited float up (bounded; zero-citation stays neutral)
search_papers(query="深度学习", language_affinity="MILD")
# demote (never exclude) results not in the query's language;
# unknown-language papers stay neutral; hard-exclude via
# filters_advanced=[{"field":"language","value":"zh"}]Search authors or journals (collection):
Set collection to authors or sources (default papers) to search
those entities. Each has its own fields — call
list_catalog(collection="authors") first; use filters_advanced +
sort_advanced (papers convenience fields apply to papers only).
search_papers(collection="authors",
filters_advanced=[{field: "summary_stats.h_index", operator: "FILTER_OP_GTE", value: 50}],
sort_advanced=[{field: "cited_by_count", order: "SORT_ORDER_DESC"}])Fetch a paper figure / image:
When read_content Markdown contains , call
get_resource with the file_name to fetch image binary.
read_content(doc_id, offset) → markdown 
└─▶ get_resource(file_name="dt=xxx/p/f3.png")Reading fulltext (check first):
Each search_papers hit carries is_content_accessible (bool): true only when
the paper has fulltext AND the caller is authorized. Check it before
read_content(doc_id, ...) — false means no fulltext or no read permission.
0 — success; stdout is the JSON response1 — HTTP 4xx/5xx; stderr contains status code and response body2 — argument error (missing token, malformed JSON, required field absent)© opendatalab, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 9 other files (scripts) in clawhub of opendatalab/Sciverse-Agent-Tools.
Open the folder on GitHubat commit 5246a81
Sciverse Academic Retrieval next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Sciverse Academic Retrieval this skillopendatalab/Sciverse-Agent-Tools | 119 | — | ~2.3k | Automated safety check: Pass | Apache-2.0 | |
| RAG Cite Sourceslyonzin/knowledge-rag | 292 | — | ~1.4k | Automated safety check: Pass | MIT | |
| Scholar RAGjoshzyj/open-scholar-skill | 168 | — | ~7.4k | Automated safety check: Notes | Custom licence | |
| Citation Systemkangarooking/system-prompt-skills | 207 | — | ~655 | Automated safety check: Pass | MIT | |
| Academic AioAperivue/medsci-skills | 331 | — | ~4.8k | Automated safety check: Pass | MIT | |
| Tw Legal RAGaa0101181514/tw-legal-rag | 328 | — | ~580 | Automated safety check: Pass | Custom licence |
lyonzin/knowledge-rag
Every technical claim drawn from the local corpus must ship with a source citation formatted as path:line or path:section.
joshzyj/open-scholar-skill
Build and query a local vector database + GraphRAG over your entire reference library (Zotero or a PDF folder) for literature review.
kangarooking/system-prompt-skills
当系统提示需要设计引用格式、信息溯源机制、来源标注系统时调用。适用于文档问答、搜索增强生成(RAG)、代码引用、浏览器辅助等需要让用户追溯信息来源的场景。不适用于纯创作类输出(如故事、诗歌),不适用于无需溯源的常识问答,也不适用于注入防御(虽然两者都涉及内容可信度)。
Aperivue/medsci-skills
A skill your agent uses when a medical AI paper should be found and cited by AI search engines and RAG tools.
aa0101181514/tw-legal-rag
Retrieve real Taiwan court judgments with verifiable citations before answering any question about Taiwan law or case law.
brightdata/skills
Produce a deep, multi-source, cited research brief on a topic from live web data using Bright Data's Discover API (intent-ranked web search + parsed page content).
opendatalab/Sciverse-Agent-Tools
A skill your agent uses when the user needs academic paper retrieval — searching scientific literature by author/year/journal, finding paper chunks for RAG-style citations, or expanding original…
Categories
Sciverse academic paper retrieval: structured metadata search, semantic chunk retrieval for RAG, and character-range content reading (offsets in Unicode code points). Sciverse Academic Retrieval is an agent skill from opendatalab/Sciverse-Agent-Tools. Sciverse academic paper retrieval: structured metadata search, semantic chunk retrieval for RAG, and character-range content reading (offsets in Unicode code points).
Sciverse Academic Retrieval fits situations like: tasks that involve Citation management; tasks that involve Retrieval-augmented generation.
Run `npx skills add opendatalab/Sciverse-Agent-Tools --skill sciverse-academic-retrieval -a claude-code`. Or copy the skill folder (clawhub in opendatalab/Sciverse-Agent-Tools) into .claude/skills/sciverse-academic-retrieval in your project. Claude Code loads it when a task matches its description.
Run `npx skills add opendatalab/Sciverse-Agent-Tools --skill sciverse-academic-retrieval -a codex`. Or copy the skill folder (clawhub in opendatalab/Sciverse-Agent-Tools) into .agents/skills/sciverse-academic-retrieval in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add opendatalab/Sciverse-Agent-Tools --skill sciverse-academic-retrieval -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/sciverse-academic-retrieval, .gemini/skills/sciverse-academic-retrieval, .github/skills/sciverse-academic-retrieval and .opencode/skills/sciverse-academic-retrieval in your project.
Going by SKILL.md and its folder, Sciverse Academic Retrieval needs JavaScript for the scripts in its folder, the command-line tools its instructions call (node) and credentials named SCIVERSE_API_TOKEN. Our summary lists: Python 3; Node.js; A credential in SCIVERSE_API_TOKEN.
SKILL.md names 1 domain. As links in the text: sciverse.space. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Sciverse Academic Retrieval is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.3k tokens (SKILL.md is roughly 9.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Sciverse Academic Retrieval: RAG Cite Sources (lyonzin/knowledge-rag, 292 stars), Scholar RAG (joshzyj/open-scholar-skill, 168 stars), Citation System (kangarooking/system-prompt-skills, 207 stars) and Academic Aio (Aperivue/medsci-skills, 331 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
opendatalab (a GitHub organization) maintains it in opendatalab/Sciverse-Agent-Tools, which has 119 GitHub stars. The repository holds 2 skills in this directory. The repository was last updated on September 20, 2026.
Source: opendatalab/Sciverse-Agent-Tools on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.