Agent skill

Ingest

by ZimoLiao in ZimoLiao/scholaraio

A skill your agent uses when the user wants to process new papers, patents, theses, documents, or proceedings from inbox queues into the knowledge base, run the ingest pipeline, or rebuild indexes.

MITAuto-check passedDocuments & Office

Install Ingest

skills CLI
$ npx skills add ZimoLiao/scholaraio --skill ingest -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ZimoLiao/scholaraio ingest --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ZimoLiao/scholaraio.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/ingest .claude/skills/ingest && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ingest
GitHub stars
577
Token cost
~1.3k tokens
SKILL.md length
428 words
Files
1
Skills in repo
43
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when the user wants to process new papers, patents, theses, documents, or proceedings from inbox queues into the knowledge base, run the ingest pipeline, or rebuild indexes.

  • Works in 2 steps: 根据用户意图选择预设 → 执行流水线命令
  • The user wants to process new papers
  • SKILL.md covers 支持的文件格式, 执行逻辑 and 示例
  • Calls pip

What it does

Ingest is an agent skill from ZimoLiao/scholaraio. Use when the user wants to process new papers, patents, theses, documents, or proceedings from inbox queues into the knowledge base, run the ingest pipeline, or rebuild indexes.

Its SKILL.md is about 1.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Documents & Office, covering Email management, Intellectual property and Knowledge bases. It works with Microsoft Excel. The repository describes itself as: Scholar All-In-One: A research infrastructure for AI agents. The licence is MIT.

When your agent uses it

  • The user wants to process new papers
  • Proceedings from inbox queues into the knowledge base
  • Run the ingest pipeline
  • Rebuild indexes

Example prompts

  • “/ingest”

Requirements

  • Python 3

Workflow steps

2 steps, taken from the first numbered list in SKILL.md.

  1. 根据用户意图选择预设
  2. 执行流水线命令

What it can do on your machine

Read from SKILL.md and the folder at commit 777628b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Ingest loads about 1.3k tokens when it runs. Until then it costs about 46 tokens; SKILL.md has 428 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~46
When it runs · the whole SKILL.md, loaded when a task matches
~1.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from ZimoLiao/scholaraio at commit 777628b, republished under its MIT licence (© ZimoLiao). 428 words, ~1,330 tokens.

Download SKILL.mdSave it as .claude/skills/ingest/SKILL.md (or your agent's skills folder).
name
ingest
description
Use when the user wants to process new papers, patents, theses, documents, or proceedings from inbox queues into the knowledge base, run the ingest pipeline, or rebuild indexes.

入库文档

将 inbox 中的 PDF、Office 文档(DOCX/XLSX/PPTX)或 Markdown 文件处理入库。支持论文、专利、学位论文、一般文档和论文集(proceedings)。

支持的文件格式

格式放入目录处理方式
.pdfdata/spool/inbox/ 或 data/spool/inbox-doc/MinerU 转 Markdown
.pdf / .mddata/spool/inbox-patent/专利文献(按公开号去重)
.pdf / .mddata/spool/inbox-proceedings/论文集准备流程(先生成 proceeding.md + split_candidates.json)
.docx .xlsx .pptxdata/spool/inbox-doc/MarkItDown 转 Markdown
.md任意 inbox直接入库(跳过转换)

执行逻辑

  1. 根据用户意图选择预设:

    • 入库新文档(默认):使用 ingest 预设(= mineru, extract, dedup, ingest, embed, index)
    • 完整处理:使用 full 预设(= mineru, extract, dedup, ingest, toc, l3, embed, index)
    • 仅重建索引:使用 reindex 预设(= embed, index)
    • 仅内容富化:使用 enrich 预设(= toc, l3, embed, index)

    注意:inbox-doc/ 始终使用专用步骤 office_convert, mineru, extract_doc, ingest,不受 preset 影响。inbox-patent/ 和 inbox-thesis/ 也有各自的固定流程。preset 中的 papers 级步骤(toc, l3)和 global 级步骤(embed, index)在处理完所有 inbox 后统一执行。

  2. 执行流水线命令:

bash
scholaraio pipeline <preset> [--dry-run] [--no-api] [--force] [--inspect]

可用预设:full | ingest | enrich | reindex

常用选项:

  • --dry-run — 预览处理,不写文件
  • --no-api — 离线模式,跳过外部 API 查询
  • --force — 强制重新处理(toc/l3 等步骤)
  • --inspect — 展示处理详情
  • --steps STEPS — 自定义步骤序列(逗号分隔),如 --steps toc,l3,index
  • --list — 列出所有可用步骤和预设
  1. pipeline 当前会依次处理五个 inbox 目录:

    • data/spool/inbox/ — 普通论文(有 DOI 才入库,无 DOI 且非 thesis 转 pending)
    • data/spool/inbox-thesis/ — 学位论文(跳过 DOI 去重,自动标记 thesis)
    • data/spool/inbox-patent/ — 专利文献(按公开号去重,自动标记 patent,跳过 DOI 去重)
    • data/spool/inbox-doc/ — 非论文文档(技术报告、讲义、Word/Excel/PPT、标准文档等,跳过 DOI 去重,LLM 生成标题/摘要)
    • data/spool/inbox-proceedings/ — 论文集(强制按 proceedings 处理;普通 data/spool/inbox/ 不会当作 proceedings)

    旧版 data/inbox* 与 data/pending/ 是迁移输入,不是当前正常 runtime 输入。先运行 scholaraio migrate upgrade --migration-id <id> --confirm,再执行入库流程。

  2. 论文类的 Stage-1 元数据提取由 ingest.extractor 控制:

    • regex:纯正则,最快,不调用 LLM
    • auto:正则优先,关键字段缺失时再调用 LLM
    • robust:正则 + LLM 双跑,校正 OCR 错误并处理多 DOI 情况(默认)
    • llm:纯 LLM 提取
    • 如果用户问“为什么标题 / 作者 / DOI 提取不准”,先检查这里的模式配置
  3. 论文集(proceedings)采用半自动两阶段流程:

    • 第一阶段:scholaraio pipeline ingest 只负责把 PDF/MD 转成 configured proceedings library(fresh 默认 data/libraries/proceedings/<Volume>/proceeding.md),并生成 split_candidates.json
    • 此时不会自动拆成子论文;CLI 会显式提示等待 agent 审阅 split_candidates.json 并生成 split_plan.json
    • 第二阶段:由 agent/人工审阅结构后,执行
bash
scholaraio proceedings apply-split <proceeding_dir> <split_plan.json>
  • 这一步才会真正把子论文落到 configured proceedings library 的 <Volume>/papers/<Paper>/
  1. proceedings 拆分后支持半自动清洗流程:
    • 先执行
bash
scholaraio proceedings build-clean-candidates <proceeding_dir>
  • 该命令会生成 clean_candidates.json,用于汇总每个 child paper 的开头窗口、heading、缺失字段和结构信号
  • 然后由 agent/人工审阅并生成 clean_plan.json
  • 最后执行
bash
scholaraio proceedings apply-clean <proceeding_dir> <clean_plan.json>
  • 第一版支持的清洗动作是 keep / rename / reclassify / drop
  • agent 在这一步还可以顺手删除明显不合理的标签行,例如假 # Comment 2.、假 # Reporter ...
  • 这里的“删除标签”只针对明显错误的独立 heading/tag 行,不改正文段落内容
  • 推荐先做结构性清洗(保留/重命名/重分类/删除),再考虑作者、摘要、DOI 等元数据提纯
  1. Office 文件处理流程(data/spool/inbox-doc/ 中的 DOCX/XLSX/PPTX):

    • step_office_convert(MarkItDown)→ 转换为 <stem>.md
    • step_extract_doc(LLM 生成标题/摘要)
    • step_ingest(写入 configured papers library,fresh 默认 data/libraries/papers/)
    • 依赖:需安装 pip install 'markitdown[docx,pptx,xlsx]'
  2. 专利文献处理逻辑(data/spool/inbox-patent/):

    • 自动提取公开号(CN/US/EP/WO/JP/KR/DE/FR/GB/TW/IN/AU 等格式)
    • 按公开号去重(非 DOI),跳过 DOI 检查
    • 自动标记 paper_type: patent
  3. 无 DOI 论文的处理逻辑:

    • 来自 data/spool/inbox-thesis/ → 直接标记为 thesis 并入库
    • 来自 data/spool/inbox-doc/ → 标记为 document 类型,LLM 生成标题和摘要后入库
    • 来自 data/spool/inbox/ → LLM 分析判断是否 thesis
      • 是 thesis → 标记并入库
      • 不是 thesis → 转入 data/spool/pending/ 待人工确认
  4. 超长 PDF 会在 MinerU 转换前按需自动切分后合并:

Show full SKILL.md (113 more words)Show less
  • 本地 MinerU 按 chunk_page_limit(默认 >100 页)
  • 云端 MinerU 同时遵循 >600 页 和 >200MB 两个限制,并在仅超大小时估算更安全的分片页数
  1. 如果 config.translate.auto_translate: true,只要本次 pipeline 包含 inbox 步骤并成功入库新论文,系统会在 papers 阶段自动插入 translate,位置在 embed/index 之前:
  • 只翻译本次新入库论文,不会顺手重翻整个库
  • 目标语言读取 translate.target_lang
  • 这是配置驱动行为,不需要额外改 preset

示例

用户说:"我放了几篇新论文到 inbox,帮我入库" → 执行 pipeline ingest

用户说:"把这个网页/在线 PDF 直接收进库里" → 先用当前 Agent 的原生网页读取能力获取内容;只有用户明确要求持久化时,才把可审查的正文保存为受支持的 inbox 文件并走正常 ingest 流程

用户说:"我在校园网/机构网,帮我从 DOI 或出版社页面下载正版论文 PDF" → 使用 scholaraio fetch-pdf <doi-or-url> --direct;如果要马上入库,加 --ingest。这个命令只利用用户当前合法访问上下文,不做访问绕过,也不需要 Paper Fetch Skill 或 PDF 转换功能。

用户说:"把新论文全部处理完,包括提取目录和结论" → 执行 pipeline full

用户说:"我有几份技术报告放在 inbox-doc 里了" → 执行 pipeline ingest(pipeline 自动处理五个 inbox 目录)

用户说:"我把一个 Word 文档放进 inbox-doc 了" → 执行 pipeline ingest(自动用 MarkItDown 转换 DOCX)

用户说:"我有几篇专利放在 inbox-patent 了" → 执行 pipeline ingest(自动处理五个 inbox 目录,专利按公开号去重)

用户说:"我有一本文集放在 inbox-proceedings 里" → 先执行 pipeline ingest,等生成 split_candidates.json 后由 agent 审阅,再执行 scholaraio proceedings apply-split ...

用户说:"重新建索引" → 执行 pipeline reindex

© ZimoLiao, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/ingest of ZimoLiao/scholaraio.

Open the folder on GitHubat commit 777628b

Compare with similar skills

Ingest next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Ingest compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Ingest this skillZimoLiao/scholaraio577—~1.3kAutomated safety check: PassMIT
Local Knowledge Base RetrieverConardLi/rag-skill7152 repos~1.7kAutomated safety check: PassNone
aai-cli Microsoft 365aai-labs/agent-barn109—~1.2kAutomated safety check: PassApache-2.0
Tencent DocsLeoYeAI/openclaw-master-skills2.2k—~3kAutomated safety check: PassMIT
Knowledge Ingestevolution-foundation/evo-nexus545—~922Automated safety check: PassCustom licence
Paper2patent7toCR/paper2patent654—~2.5kAutomated safety check: PassMIT

Similar skills

  • Answers questions from a local knowledge base folder by walking hierarchical index files and searching with grep, pdfplumber and pandas instead of loading whole files.

    715 GitHub starsUsed in 2 repos~1.7k tokens
    Knowledge ManagementAuto-check passed
  • aai-cli Microsoft 365

    aai-labs/agent-barn

    Guides work with Outlook, OneDrive, SharePoint, Teams, Excel, To Do and Planner through aai-cli's Microsoft Graph commands, starting from which service owns the data.

    109 GitHub stars~1.2k tokensUpdated today
    Documents & OfficeAuto-check passed
  • Tencent Docs

    LeoYeAI/openclaw-master-skills

    Tencent Docs - Provides complete Tencent Docs operations. An agent skill from LeoYeAI/openclaw-master-skills.

    2.2k GitHub stars~3k tokensUpdated 2 mo ago
    Documents & OfficeAuto-check passed
  • Knowledge Ingest

    evolution-foundation/evo-nexus

    Upload a file (PDF, DOCX, PPTX, XLSX, HTML, EPUB, image) or URL to the Knowledge base.

    545 GitHub stars~922 tokensUpdated 4 mo ago
    Documents & OfficeAuto-check passed
  • Paper2patent

    7toCR/paper2patent

    Turn an academic paper (PDF, LaTeX, pasted text, thesis chapter or technical disclosure) into a Chinese invention patent application draft — 说明书摘要、摘要附图、权利要求书、说明书、说明书附图 — delivered as DOCX/PDF with…

    654 GitHub stars~2.5k tokensUpdated today
    Documents & OfficeAuto-check passed
  • Contract Review Engine

    infometa/workbuddyskills

    Review contracts for risk (scenarios C5/C6) — background assessment before signing (counterparty qualification, transaction-mode legality, contract-form fit, special procedures) and clause-by-clause…

    348 GitHub stars~1.8k tokensUpdated today
    Documents & OfficeAuto-check passed

More from ZimoLiao/scholaraio

All 43 skills in this repo
  • Document

    ZimoLiao/scholaraio

    A skill your agent uses when the user wants to create or inspect DOCX, PPTX, or XLSX files, generate a downloadable Office deliverable, or verify its structure and layout warnings with scholaraio…

    577 GitHub stars~674 tokensUpdated 14 days ago
    Auto-check passed
  • Academic Writing

    ZimoLiao/scholaraio

    A skill your agent uses when the user needs help choosing or organizing an academic-writing workflow by deliverable, stage, or format, including review articles, guided reading, paper sections, PPT…

    577 GitHub stars~827 tokensUpdated 14 days ago
    Auto-check passed
  • Arxiv

    ZimoLiao/scholaraio

    A skill your agent uses when the user wants to browse arXiv preprints, search arXiv directly, fetch a PDF by arXiv ID or URL, or send a preprint into the ScholarAIO ingest pipeline.

    577 GitHub stars~799 tokensUpdated 14 days ago
    Auto-check passed
  • Bioinformatics

    ZimoLiao/scholaraio

    A skill your agent uses when working on bioinformatics workflows such as alignment, variant calling, phylogenetics, or protein-structure analysis, especially across BLAST, minimap2, samtools…

    577 GitHub stars~1.4k tokensUpdated 14 days ago
    Auto-check passed
  • Citation Check

    ZimoLiao/scholaraio

    A skill your agent uses when the user wants to verify citations in AI-generated or human-written text against the local knowledge base and catch hallucinated, wrong, or missing references.

    577 GitHub stars~454 tokensUpdated 14 days ago
    Auto-check passed
  • Draw

    ZimoLiao/scholaraio

    A skill your agent uses when the user wants diagrams, flowcharts, architecture visuals, data relationships, timelines, concept maps, Mermaid, Graphviz, drawio, or polished paper figures generated…

    577 GitHub stars~1.3k tokensUpdated 14 days ago
    Auto-check: notes

Works with

Questions about Ingest

What does Ingest do?

A skill your agent uses when the user wants to process new papers, patents, theses, documents, or proceedings from inbox queues into the knowledge base, run the ingest pipeline, or rebuild indexes. Ingest is an agent skill from ZimoLiao/scholaraio. Use when the user wants to process new papers, patents, theses, documents, or proceedings from inbox queues into the knowledge base, run the ingest pipeline, or rebuild indexes.

When should I use Ingest?

Ingest fits situations like: the user wants to process new papers; proceedings from inbox queues into the knowledge base; run the ingest pipeline; rebuild indexes.

How do I install Ingest in Claude Code?

Run `npx skills add ZimoLiao/scholaraio --skill ingest -a claude-code`. Or copy the skill folder (.claude/skills/ingest in ZimoLiao/scholaraio) into .claude/skills/ingest in your project. Claude Code loads it when a task matches its description.

How do I install Ingest in Codex?

Run `npx skills add ZimoLiao/scholaraio --skill ingest -a codex`. Or copy the skill folder (.claude/skills/ingest in ZimoLiao/scholaraio) into .agents/skills/ingest in your project. Codex loads it when a task matches its description.

Can I use Ingest in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ZimoLiao/scholaraio --skill ingest -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ingest, .gemini/skills/ingest, .github/skills/ingest and .opencode/skills/ingest in your project.

What does Ingest need to run?

Going by SKILL.md and its folder, Ingest needs the command-line tools its instructions call (pip). Our summary lists: Python 3.

Does Ingest access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Ingest safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Ingest use?

Ingest is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Ingest use?

About 1.3k tokens (SKILL.md is roughly 5.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Ingest?

Skills that share tags, products or a category with Ingest: Local Knowledge Base Retriever (ConardLi/rag-skill, 715 stars), aai-cli Microsoft 365 (aai-labs/agent-barn, 109 stars), Tencent Docs (LeoYeAI/openclaw-master-skills, 2.2k stars) and Knowledge Ingest (evolution-foundation/evo-nexus, 545 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Ingest?

ZimoLiao (a GitHub user) maintains it in ZimoLiao/scholaraio, which has 577 GitHub stars. The repository holds 43 skills in this directory. The repository was last updated on September 25, 2026.

Source: ZimoLiao/scholaraio on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.