Agent skill

Sciverse Academic Retrieval

by opendatalab in opendatalab/Sciverse-Agent-Tools

Sciverse academic paper retrieval: structured metadata search, semantic chunk retrieval for RAG, and character-range content reading (offsets in Unicode code points).

Apache-2.0Auto-check passedAI & LLM Engineering

Install Sciverse Academic Retrieval

skills CLI
$ npx skills add opendatalab/Sciverse-Agent-Tools --skill sciverse-academic-retrieval -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install opendatalab/Sciverse-Agent-Tools sciverse-academic-retrieval --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/opendatalab/Sciverse-Agent-Tools.git skills-src && mkdir -p .claude/skills && cp -r skills-src/clawhub .claude/skills/sciverse-academic-retrieval && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
sciverse-academic-retrieval
GitHub stars
119
Token cost
~2.3k tokens
SKILL.md length
918 words
Files
10 (incl. scripts)
Skills in repo
2
Repo updated
First seen
Licence
Apache-2.0

At a glance

Sciverse academic paper retrieval: structured metadata search, semantic chunk retrieval for RAG, and character-range content reading (offsets in Unicode code points).

  • Tasks that involve Citation management
  • SKILL.md covers When to use, Authentication, Tools and Bootstrap: learn the schema…, plus 2 more sections
  • Runs JavaScript scripts from its folder; calls node; needs SCIVERSE_API_TOKEN
  • Tasks that involve Retrieval-augmented generation

What it does

Sciverse Academic Retrieval is an agent skill from opendatalab/Sciverse-Agent-Tools. Sciverse academic paper retrieval: structured metadata search, semantic chunk retrieval for RAG, and character-range content reading (offsets in Unicode code points). For agent workflows that need citation-grade scientific literature.

Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 10 other files, including scripts (for example `README.md` and `manifest.json`).

It sits in AI & LLM Engineering, covering Citation management and Retrieval-augmented generation. The repository describes itself as: Standardized tool schemas and SDKs that expose Sciverse Open Platform retrieval capabilities to LLM agents. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Citation management
  • Tasks that involve Retrieval-augmented generation

Example prompts

  • “/sciverse-academic-retrieval”

Requirements

  • Python 3
  • Node.js
  • A credential in SCIVERSE_API_TOKEN

What it can do on your machine

Read from SKILL.md and the folder at commit 5246a81. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 7 files in scripts/ (JavaScript), which the agent can run.

    Shell commands in SKILL.md call:

    • node

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • sciverse.space

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • SCIVERSE_API_TOKEN

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Sciverse Academic Retrieval loads about 2.3k tokens when it runs. Until then it costs about 66 tokens; SKILL.md has 918 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~66
When it runs · the whole SKILL.md, loaded when a task matches
~2.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from opendatalab/Sciverse-Agent-Tools at commit 5246a81, republished under its Apache-2.0 licence (© opendatalab). 918 words, ~2,323 tokens.

Download SKILL.mdSave it as .claude/skills/sciverse-academic-retrieval/SKILL.md (or your agent's skills folder). This skill also uses 9 other files; get the full folder from GitHub.
name
sciverse-academic-retrieval
description
Sciverse academic paper retrieval: structured metadata search, semantic chunk retrieval for RAG, and character-range content reading (offsets in Unicode code points). For agent workflows that need citation-grade scientific literature.
slug
academic-retrieval
version
0.14.3
license
Apache-2.0
homepage
https://sciverse.space

academic-retrieval

Sciverse academic paper retrieval: structured metadata search, semantic chunk retrieval for RAG, and character-range content reading (offsets in Unicode code points). For agent workflows that need citation-grade scientific literature.

When to use

Trigger this skill when the user's request involves any of:

  • Locating academic papers by structured criteria (authors, year, journal, subjects)
  • Grounding answers in paper excerpts (RAG / citations)
  • Expanding the original text around a known doc_id (more text before/after a chunk)

Authentication

This skill requires the SCIVERSE_API_TOKEN environment variable (obtain from https://sciverse.space). Optionally set SCIVERSE_BASE_URL to override the default API base URL.

Tools

search_papers

Search academic papers by structured filters (title, authors, journal, year, subjects, etc.). Use when: "find Hinton's papers from 2020-2023", "Nature papers on CRISPR". Not for: natural-language Q&A retrieval (use semantic_search) or full-text snippets (use read_content). Returns: list of papers; each entry has unique_id (always present), doc_id (only when full text exists), title, author, abstract, publication_venue_name_unified, publication_published_year.

Invoke: node scripts/search_papers.mjs '<JSON args>'

Natural-language semantic search returning relevant paper chunks for RAG-style answering. Use when: "How does Transformer attention work?", "What are recent methods for protein structure prediction?". Not for: precise field filtering (use search_papers) or fetching full original text (use read_content). Returns: list of chunks; each entry has chunk_id, doc_id, abstract, chunk, score, title, offset. Typical chain: semantic_search → pick chunk → read_content(doc_id, offset).

Invoke: node scripts/semantic_search.mjs '<JSON args>'

list_catalog

Returns the schema catalog for search_papers: every field name, type, whether it's filterable / sortable, default-return status, human description, and applicable FilterOperators. Use when: "Which field do I filter by DOI?", "What values can access_oa_status take?", "What's the right enum for metadata_type?". Not for: actually searching papers (use search_papers / semantic_search). Typical pattern: call once when first encountering Sciverse or facing an ambiguous field need, then construct precise search_papers filters from the returned schema. Pass include_sample_values=true to also fetch top-20 values for enum-like fields (OpenSearch terms aggregation, 24h cached).

Invoke: node scripts/list_catalog.mjs '<JSON args>'

list_paper_relations

Paginate the full relation list of a paper. citations/references/related_works are unbounded arrays (up to 340k entries for a single paper) and are NOT projectable in search_papers, so this endpoint is the only way to read them. Use when: "What does paper X cite?" (relation=REFERENCES), "Which papers cite paper X?" (relation=CITATIONS), "Works related to paper X" (relation=RELATED_WORKS). Note: CITATIONS (incoming: who cites me) and REFERENCES (outgoing: who I cite) are opposite directions. Typical chain: get unique_id from search_papers / semantic_search, then paginate here by relation. Two limits (CITATIONS only; REFERENCES/RELATED_WORKS max out at 11833/20 in practice): more than 10000 relations returns 429; page*page_size above 10000 returns 400. In both cases switch to search_papers with filters_advanced on references_unique_id — it supports deep paging and arbitrary sorting. total_count counts in-corpus matches only, so it can differ from the paper's own citation_count by about 1%.

Invoke: node scripts/list_paper_relations.mjs '<JSON args>'

read_content

Read a range of a paper's original text addressed in Unicode code points (offset/limit count characters like Python len(), not bytes). Typically used with a doc_id/offset returned by semantic_search to expand context (read more text before or after a chunk). Returns: text fragment, bytes_returned (UTF-8 byte length of text, for reference only), next_offset (code-point offset of the next fragment — page with it, never with bytes_returned), more (boolean). Server behaviour: limit above 524288 is silently clamped; omitting offset returns the whole document ignoring limit — the SDKs / MCP server send offset=0 and limit=4096 by default, so pass offset explicitly when calling the HTTP API directly.

Invoke: node scripts/read_content.mjs '<JSON args>'

Show full SKILL.md (348 more words)Show less
get_resource

Returns the binary bytes of a paper figure / table image referenced inside read_content's Markdown via ![alt](file_name) placeholders. Use when the user asks to see / display / describe a figure and read_content output contains an image reference. Input file_name comes from the Markdown URL part (relative path, no \\ or ..). Returns: raw image stream + image/* Content-Type. The SDK / MCP server wraps the bytes as base64 + mimeType so Claude (multimodal) can read the image directly.

Invoke: node scripts/get_resource.mjs '<JSON args>'

Bootstrap: learn the schema first

If you're unsure which fields exist or what values an enum takes (e.g. metadata_type, language, access_oa_status), call list_catalog once at the start. Sample values are returned for low-cardinality fields. Use it instead of guessing field names — guessing wastes turns.

list_catalog(include_sample_values=true)
    └─▶ fields[].name + sample_values  →  precise filter construction

Recipes

RAG flow (natural-language Q&A):

semantic_search(query=...) → hits[i].doc_id, hits[i].offset
    └─▶ read_content(doc_id, offset)

Lookup by DOI:

search_papers(filters_advanced=[{field: "doi", value: "10.1038/..."}])

OA + year filter:

search_papers(
    year_from=2024,
    filters_advanced=[{field: "access_is_oa", value: "true"}]
)

Scoped semantic search (constrained corpus):

semantic_search(
    query="...",
    filters={"author": ["Hinton"],
             "publication_published_year": {"gte": 2020}}
)   # applied at recall time, server-side; AND across fields

Soft semantics: chunks missing that metadata are NOT excluded. For a hard guarantee, or meta-only constraints (fwci, citation graph, complex hit-sets), scope by doc_id — a HARD recall-time filter:

search_papers(..., fields=["doc_id","title"]) → collect doc_id
semantic_search(query=..., filters={"doc_id": [...]})
    # hits never leave the set; empty list → empty hits (never global);
    # up to 1000 deduped ids (400 SCOPE_TOO_LARGE beyond)

Bias fuzzy search ranking (soft boosts — stackable):

Three multiplicative boosts (freshness_boost / impact_boost / language_affinity, each NONE/MILD/STRONG) reorder fuzzy-search results while keeping relevance. Only effective when query is non-empty; ignored when any sort is set; shallow paging while active. sort_by_year defaults to auto (relevance with query, newest-first for pure filters); query+desc is an anti-pattern — it degrades the query to a match filter and disables all boosts; use freshness_boost.

search_papers(query="large language model", freshness_boost="STRONG")
    # recent first: STRONG=3-year decay, MILD=10-year
search_papers(query="protein folding", impact_boost="MILD")
    # highly-cited float up (bounded; zero-citation stays neutral)
search_papers(query="深度学习", language_affinity="MILD")
    # demote (never exclude) results not in the query's language;
    # unknown-language papers stay neutral; hard-exclude via
    # filters_advanced=[{"field":"language","value":"zh"}]

Search authors or journals (collection):

Set collection to authors or sources (default papers) to search those entities. Each has its own fields — call list_catalog(collection="authors") first; use filters_advanced + sort_advanced (papers convenience fields apply to papers only).

search_papers(collection="authors",
    filters_advanced=[{field: "summary_stats.h_index", operator: "FILTER_OP_GTE", value: 50}],
    sort_advanced=[{field: "cited_by_count", order: "SORT_ORDER_DESC"}])

Fetch a paper figure / image:

When read_content Markdown contains ![alt](file_name), call get_resource with the file_name to fetch image binary.

read_content(doc_id, offset) → markdown ![Figure 3](dt=xxx/p/f3.png)
    └─▶ get_resource(file_name="dt=xxx/p/f3.png")

Reading fulltext (check first):

Each search_papers hit carries is_content_accessible (bool): true only when the paper has fulltext AND the caller is authorized. Check it before read_content(doc_id, ...) — false means no fulltext or no read permission.

Exit codes

  • 0 — success; stdout is the JSON response
  • 1 — HTTP 4xx/5xx; stderr contains status code and response body
  • 2 — argument error (missing token, malformed JSON, required field absent)

© opendatalab, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 9 other files (scripts) in clawhub of opendatalab/Sciverse-Agent-Tools.

  • SKILL.md
  • README.md
  • manifest.json
  • scripts/_common.mjs
  • scripts/get_resource.mjs
  • scripts/list_catalog.mjs
  • scripts/list_paper_relations.mjs
  • scripts/read_content.mjs
  • scripts/search_papers.mjs
  • scripts/semantic_search.mjs

Open the folder on GitHubat commit 5246a81

Compare with similar skills

Sciverse Academic Retrieval next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Sciverse Academic Retrieval compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Sciverse Academic Retrieval this skillopendatalab/Sciverse-Agent-Tools119—~2.3kAutomated safety check: PassApache-2.0
RAG Cite Sourceslyonzin/knowledge-rag292—~1.4kAutomated safety check: PassMIT
Scholar RAGjoshzyj/open-scholar-skill168—~7.4kAutomated safety check: NotesCustom licence
Citation Systemkangarooking/system-prompt-skills207—~655Automated safety check: PassMIT
Academic AioAperivue/medsci-skills331—~4.8kAutomated safety check: PassMIT
Tw Legal RAGaa0101181514/tw-legal-rag328—~580Automated safety check: PassCustom licence

Similar skills

  • RAG Cite Sources

    lyonzin/knowledge-rag

    Every technical claim drawn from the local corpus must ship with a source citation formatted as path:line or path:section.

    292 GitHub stars~1.4k tokensUpdated 5 days ago
    AI & LLM EngineeringAuto-check passed
  • Scholar RAG

    joshzyj/open-scholar-skill

    Build and query a local vector database + GraphRAG over your entire reference library (Zotero or a PDF folder) for literature review.

    168 GitHub stars~7.4k tokensUpdated 21 days ago
    Research & ScienceAuto-check: notes
  • Citation System

    kangarooking/system-prompt-skills

    当系统提示需要设计引用格式、信息溯源机制、来源标注系统时调用。适用于文档问答、搜索增强生成(RAG)、代码引用、浏览器辅助等需要让用户追溯信息来源的场景。不适用于纯创作类输出(如故事、诗歌),不适用于无需溯源的常识问答,也不适用于注入防御(虽然两者都涉及内容可信度)。

    207 GitHub stars~655 tokensUpdated 5 mo ago
    AI & LLM EngineeringAuto-check passed
  • Academic Aio

    Aperivue/medsci-skills

    A skill your agent uses when a medical AI paper should be found and cited by AI search engines and RAG tools.

    331 GitHub stars~4.8k tokensUpdated 4 days ago
    Research & ScienceAuto-check passed
  • Tw Legal RAG

    aa0101181514/tw-legal-rag

    Retrieve real Taiwan court judgments with verifiable citations before answering any question about Taiwan law or case law.

    328 GitHub stars~580 tokensUpdated yesterday
    Legal & ComplianceAuto-check passed
  • Live Research

    brightdata/skills

    Produce a deep, multi-source, cited research brief on a topic from live web data using Bright Data's Discover API (intent-ranked web search + parsed page content).

    264 GitHub stars~1.8k tokensUpdated 2 days ago
    Research & ScienceAuto-check passed

More from opendatalab/Sciverse-Agent-Tools

  • Sciverse

    opendatalab/Sciverse-Agent-Tools

    A skill your agent uses when the user needs academic paper retrieval — searching scientific literature by author/year/journal, finding paper chunks for RAG-style citations, or expanding original…

    119 GitHub stars~3k tokensUpdated 19 days ago
    Auto-check passed

Questions about Sciverse Academic Retrieval

What does Sciverse Academic Retrieval do?

Sciverse academic paper retrieval: structured metadata search, semantic chunk retrieval for RAG, and character-range content reading (offsets in Unicode code points). Sciverse Academic Retrieval is an agent skill from opendatalab/Sciverse-Agent-Tools. Sciverse academic paper retrieval: structured metadata search, semantic chunk retrieval for RAG, and character-range content reading (offsets in Unicode code points).

When should I use Sciverse Academic Retrieval?

Sciverse Academic Retrieval fits situations like: tasks that involve Citation management; tasks that involve Retrieval-augmented generation.

How do I install Sciverse Academic Retrieval in Claude Code?

Run `npx skills add opendatalab/Sciverse-Agent-Tools --skill sciverse-academic-retrieval -a claude-code`. Or copy the skill folder (clawhub in opendatalab/Sciverse-Agent-Tools) into .claude/skills/sciverse-academic-retrieval in your project. Claude Code loads it when a task matches its description.

How do I install Sciverse Academic Retrieval in Codex?

Run `npx skills add opendatalab/Sciverse-Agent-Tools --skill sciverse-academic-retrieval -a codex`. Or copy the skill folder (clawhub in opendatalab/Sciverse-Agent-Tools) into .agents/skills/sciverse-academic-retrieval in your project. Codex loads it when a task matches its description.

Can I use Sciverse Academic Retrieval in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add opendatalab/Sciverse-Agent-Tools --skill sciverse-academic-retrieval -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/sciverse-academic-retrieval, .gemini/skills/sciverse-academic-retrieval, .github/skills/sciverse-academic-retrieval and .opencode/skills/sciverse-academic-retrieval in your project.

What does Sciverse Academic Retrieval need to run?

Going by SKILL.md and its folder, Sciverse Academic Retrieval needs JavaScript for the scripts in its folder, the command-line tools its instructions call (node) and credentials named SCIVERSE_API_TOKEN. Our summary lists: Python 3; Node.js; A credential in SCIVERSE_API_TOKEN.

Does Sciverse Academic Retrieval access the network?

SKILL.md names 1 domain. As links in the text: sciverse.space. This is read from the text; nothing was executed.

Is Sciverse Academic Retrieval safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Sciverse Academic Retrieval use?

Sciverse Academic Retrieval is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Sciverse Academic Retrieval use?

About 2.3k tokens (SKILL.md is roughly 9.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Sciverse Academic Retrieval?

Skills that share tags, products or a category with Sciverse Academic Retrieval: RAG Cite Sources (lyonzin/knowledge-rag, 292 stars), Scholar RAG (joshzyj/open-scholar-skill, 168 stars), Citation System (kangarooking/system-prompt-skills, 207 stars) and Academic Aio (Aperivue/medsci-skills, 331 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Sciverse Academic Retrieval?

opendatalab (a GitHub organization) maintains it in opendatalab/Sciverse-Agent-Tools, which has 119 GitHub stars. The repository holds 2 skills in this directory. The repository was last updated on September 20, 2026.

Source: opendatalab/Sciverse-Agent-Tools on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.