Agent skill

Doc Importer

by wpsnote in wpsnote/wpsnote-skills

将本地文档批量导入到 WPS 笔记。支持扫描 Obsidian Vault、思源笔记、微信公众号存档、 下载目录或任意用户指定目录,自动识别 HTML、Markdown、PDF、DOCX、PPTX、XLSX 等格式, 转换为 WPS 笔记可读内容并保留图片和富文本格式。

MITAuto-check passedDocuments & Office

Install Doc Importer

skills CLI
$ npx skills add wpsnote/wpsnote-skills --skill doc-importer -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install wpsnote/wpsnote-skills doc-importer --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/wpsnote/wpsnote-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/doc-importer .claude/skills/doc-importer && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
doc-importer
GitHub stars
179
Token cost
~2.7k tokens
SKILL.md length
454 words
Files
4 (incl. scripts)
Skills in repo
37
Repo updated
First seen
Licence
MIT

At a glance

将本地文档批量导入到 WPS 笔记。支持扫描 Obsidian Vault、思源笔记、微信公众号存档、 下载目录或任意用户指定目录,自动识别 HTML、Markdown、PDF、DOCX、PPTX、XLSX 等格式, 转换为 WPS 笔记可读内容并保留图片和富文本格式。

  • Works in 5 steps: get_note_outline 默认只返回 100 个 block → insert_image 在旧版 WPS 中只能插到前台笔记 → list_notes 返回大数据时要主动收敛范围 → …
  • Tasks that involve PowerPoint presentations
  • SKILL.md covers 快速开始, 支持的文档来源, 支持的文件格式 and WPS API 关键限制(必读), plus 5 more sections
  • Runs Python scripts from its folder; calls pip3 and brew

What it does

Doc Importer is an agent skill from wpsnote/wpsnote-skills. 将本地文档批量导入到 WPS 笔记。支持扫描 Obsidian Vault、思源笔记、微信公众号存档、 下载目录或任意用户指定目录,自动识别 HTML、Markdown、PDF、DOCX、PPTX、XLSX 等格式, 转换为 WPS 笔记可读内容并保留图片和富文本格式。 当用户说「导入文档到 WPS 笔记」「把我的 Obsidian 笔记导入」「导入思源笔记」 「把下载的 PDF/Word/PPT 导入笔记」「把公众号文章导入笔记」「同步本地文档到 WPS 笔记」时触发。 不适用于直接编辑 WPS 笔记内容、文档格式转换(不导入到笔记)。

Its SKILL.md is about 2.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including scripts (for example `scripts/convert.py`, `scripts/import_to_wps.py` and `scripts/scan_docs.py`).

It sits in Documents & Office, covering PowerPoint presentations, Word documents and Excel spreadsheets. It works with Obsidian, Microsoft Excel, Microsoft PowerPoint and Microsoft Word. The licence is MIT.

When your agent uses it

  • Tasks that involve PowerPoint presentations
  • Tasks that involve Word documents
  • Tasks that involve Excel spreadsheets

Example prompts

  • “/doc-importer”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. get_note_outline 默认只返回 100 个 block
  2. insert_image 在旧版 WPS 中只能插到前台笔记
  3. list_notes 返回大数据时要主动收敛范围
  4. anchor 失效导致内容静默丢失
  5. 批量写入的稳定参数

What it can do on your machine

Read from SKILL.md and the folder at commit 2d9885d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 3 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • pip3
    • brew

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip3, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Doc Importer loads about 2.7k tokens when it runs. Until then it costs about 71 tokens; SKILL.md has 454 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~71
When it runs · the whole SKILL.md, loaded when a task matches
~2.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from wpsnote/wpsnote-skills at commit 2d9885d, republished under its MIT licence (© wpsnote). 454 words, ~2,683 tokens.

Download SKILL.mdSave it as .claude/skills/doc-importer/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
doc-importer
description
将本地文档批量导入到 WPS 笔记。支持扫描 Obsidian Vault、思源笔记、微信公众号存档、 下载目录或任意用户指定目录,自动识别 HTML、Markdown、PDF、DOCX、PPTX、XLSX 等格式, 转换为 WPS 笔记可读内容并保留图片和富文本格式。 当用户说「导入文档到 WPS 笔记」「把我的 Obsidian 笔记导入」「导入思源笔记」 「把下载的 PDF/Word/PPT 导入笔记」「把公众号文章导入笔记」「同步本地文档到 WPS 笔记」时触发。 不适用于直接编辑 WPS 笔记内容、文档格式转换(不导入到笔记)。
license
MIT
metadata.author
洛小山 (itshen)
metadata.version
1.2.0
metadata.category
productivity
metadata.tags
obsidian, siyuan, wechat, html, import, pdf, docx, pptx, markdown

文档导入器(Doc Importer)

将本地文档(Obsidian、思源笔记、微信公众号 HTML、下载目录或任意目录)批量导入到 WPS 笔记。


快速开始

当前技能的笔记写入统一走笔记 Agent 工具;如需理解批量扫描和转换逻辑,可先读取本 Skill 自带脚本:

text
参考实现:
- read_file("scripts/scan_docs.py")        # 目录扫描与来源识别
- read_file("scripts/import_to_wps.py")    # 转换与导入主流程

实际写入:
- create_note
- get_note_outline
- batch_edit
- insert_image
- read_note / read_blocks
- search_notes
- sync_note

支持的文档来源

来源说明典型目录结构
Obsidian Vault扫描 .md 文件,保留 wiki 链接、Callout、标签Vault/笔记名.md + attachments/
思源笔记扫描 .sy 文件(JSON格式),提取 Block TreeSiYuan/data/笔记本/文档.sy
微信公众号解析 原文.html(含富文本格式、内联样式)文章名/原文.html + 图片/
任意目录用户指定路径,递归扫描子目录任意

支持的文件格式

格式转换方式图片处理富文本保留
.htmlBeautifulSoup 解析内联样式本地图片 base64✅ 颜色/粗体/标题
.md / .markdown直接解析,转 WPS XML本地图片 base64基本格式
.pdfpdfplumber 提取文本 + pdfimages 提取图片提取嵌入图片标题推断
.docxpandoc 转 markdown,提取 word/media/解包提取基本格式
.pptxmarkitdown 提取文本 + 解包媒体解包提取幻灯片结构
.xlsxpandas 读取,转 WPS table不含图片表格结构
.txt直接读取不含图片无
.syJSON 解析思源 Block Tree提取 assets 图片全部

WPS API 关键限制(必读)

在实际导入前,必须了解以下 WPS 笔记 API 的重要限制,否则会导致内容丢失或图片插入失败。

1. get_note_outline 默认只返回 100 个 block

get_note_outline 默认按分页视图返回 block,blocks 数组可能只覆盖前一部分内容,但 block_count 会返回真实总数。

影响:大文章(> 100 个 block)如果只靠首屏大纲查占位符,只能找到前一部分 block,后半段图片会丢失。

解决方案:先用 get_note_outline 拿首批 block,再用 read_blocks 续读:

python
def get_all_blocks(note_id):
    """翻页获取笔记全部 blocks"""
    # 第一页用 get_note_outline(有 preview 字段)
    r = get_note_outline(note_id=note_id)
    data = r.get('data', {})
    total = data.get('block_count', 0)
    blocks = list(data.get('blocks', []))
    last_id = blocks[-1]['id'] if blocks else None

    # 超过首批范围时用 read_blocks 续读
    while len(blocks) < total and last_id:
        r2 = read_blocks(note_id=note_id, block_id=last_id, after=100)
        new_blocks = (r2.get('data') or {}).get('blocks', [])
        if not new_blocks:
            break
        blocks.extend(new_blocks)
        last_id = new_blocks[-1]['id']

    return blocks

get_note_outline 返回的 block 常带 preview 字段;read_blocks 更适合拿完整 XML 内容。


2. insert_image 在旧版 WPS 中只能插到前台笔记

旧版本(< 0.1.4):insert_image 的图片 fileID 绑定到当前 WPS 客户端 UI 中打开的笔记,即使传了 note_id 参数,图片也会错误地关联到前台笔记。

新版本(>= 0.1.4):已修复,insert_image 可以直接后台插图到任意笔记,无需切换前台。

处理策略:

  • 先按后台插图能力执行;如果图片插错位置,再回退为“切到目标笔记后插图”
  • 如果宿主或客户端仍表现为旧版行为,明确提示用户升级或手动切到目标笔记
  • 插图后务必回读验收,不要假设成功

3. list_notes 返回大数据时要主动收敛范围

list_notes 一次返回的数据过大时,不适合拿来做全文去重或全量扫描判断。

解决方案:

  • 去重优先用 search_notes 按标题关键词搜索
  • 浏览列表时主动分页或缩小筛选范围
  • 不把一次大范围 list_notes 当成唯一真相来源

4. anchor 失效导致内容静默丢失

连续调用 batch_edit 插入内容时,WPS 内部可能重新索引,导致前一次操作返回的 anchor_id 失效。如果不处理,后续内容会静默跳过,造成文章后半段内容丢失。

必须实现重试逻辑:

python
def do_insert(note_id, content, get_anchor_fn, max_retries=4):
    """带重试的内容插入,自动刷新 anchor"""
    anchor = get_anchor_fn()
    for attempt in range(max_retries):
        if attempt > 0:
            time.sleep(1.5 * attempt)
            anchor = get_anchor_fn()  # 重新获取最新 anchor
        res = batch_edit(note_id, [{
            'op': 'insert', 'anchor_id': anchor,
            'position': 'after', 'content': content
        }])
        if res.get('ok') is not False:
            anchor = get_anchor_fn()
            return True
    return False  # 4次都失败才放弃

获取最新 anchor(最后一个 block):

python
def get_last_block_id(note_id):
    r = get_note_outline(note_id=note_id)
    blocks = (r.get('data') or {}).get('blocks', [])
    return blocks[-1]['id'] if blocks else None

5. 批量写入的稳定参数

大量内容写入时,BATCH_SIZE(每次 batch_edit 的 block 数)建议设为 4,过大会增加 anchor 失效概率。

python
BATCH_SIZE = 4  # 不要超过 8,否则容易出现 anchor 失效

完整工作流

第一步:确定扫描目录

优先根据用户已给出的目录 / 文件线索判断来源类型,不足时再 ask_user 补最少必要信息:

text
优先识别这些常见来源线索:
- Obsidian Vault:`.md` 文件 + `attachments/` 或同类附件目录
- 思源笔记:`.sy` 文件 + `data/assets/`
- 微信公众号存档:`原文.html` + 图片目录 + `meta.json`
- 下载目录混合导入:PDF / DOCX / PPTX / XLSX / Markdown / HTML 混合存在

第二步:扫描文档列表
text
如需对照扫描规则,用 `read_file("scripts/scan_docs.py")` 查看来源识别和输出字段约定;
整理后的清单至少要包含路径、标题、时间、格式、图片数和预计 block 数。

输出 JSON 格式:

json
{
  "source_type": "wechat_mp",
  "root_path": "/Users/xxx/Documents/articles",
  "files": [
    {
      "path": "/Users/xxx/Documents/articles/文章名/原文.html",
      "rel_path": "文章名/原文.html",
      "title": "文章标题",
      "publish_time": "2025-04-21 18:30",
      "size_bytes": 204800,
      "modified": "2025-04-22T10:00:00",
      "format": "html",
      "estimated_images": 12,
      "estimated_blocks": 180
    }
  ],
  "total": 71,
  "formats": {"html": 71}
}

estimated_blocks 帮助预判是否会超过 100 block,提前提示用户。


第三步:展示文件清单,询问选择
扫描到 71 个文件:
  - HTML: 71 个

文件列表:
 1. AutoGLM 发布之后,如今国产大模型终于长出了手。  (2025-03-31, 12张图)
 2. 你可能看不懂扣子空间为什么重要…                 (2025-04-21, 17张图)
 ...(超过20个时截断,告知总数)

请问你想如何导入?
 [A] 全部导入(71个文件)
 [B] 手动选择(输入文件编号,如:1,3,5-10)
 [S] 跳过已有标题的笔记(根据笔记标题去重)

第四步:去重检测

导入前通过标题检查 WPS 笔记中是否已存在同名笔记。

注意:不要把 list_notes 当成去重主入口,改用 search_notes 按标题关键词搜索:

python
def check_exists(title):
    r = search_notes(keyword=title[:20], limit=5)
    notes = (r.get('data') or {}).get('notes', [])
    return next((n for n in notes if n['title'] == title), None)

发现重复时询问用户:

发现以下笔记在 WPS 中已存在:
 - 《AutoGLM 发布之后…》(最后更新:2025-05-01)

如何处理?
 [O] 覆盖  [S] 跳过  [A] 追加  [RA] 对所有冲突应用相同策略

第五步:转换并导入

推荐流程(以 HTML 富文本为例):

python
# 1. 解析 HTML,提取内容段落和图片
segments = html_to_segments(html_path, img_dir)
# segments 格式:[('xml', '<p>文字</p>'), ('img', Path('图片/image_001.jpg')), ...]

# 2. 创建笔记
note_id = create_note(title)

# 3. 写入标题行 + meta 行(时间、标签)
write_header(note_id, title, publish_time, tag)

# 4. 批量写入正文(图片先插占位符)
write_content_with_placeholders(note_id, segments)

# 5. 翻页查找所有占位符 block_id(用 get_all_blocks)
ph_map = find_placeholders(note_id)

# 6. 逐个插入真实图片,替换占位符
for idx, img_path in img_list:
    insert_image(note_id, ph_map[idx]['block_id'], img_path)
    delete_placeholder(note_id, ph_map[idx]['block_id'])

Meta 行格式(根据来源类型灵活调整):

python
# 微信公众号:简洁的时间 + 标签
f'<p>{publish_time} | <tag id="{tag_id}">#推文</tag></p>'

# 通用文档:完整 meta blockquote
"""
<blockquote>
  <p>📄 <strong>来源</strong>:{rel_path}</p>
  <p>🕒 <strong>修改时间</strong>:{modified_time}</p>
  <p>🔄 <strong>导入时间</strong>:{import_time}</p>
</blockquote>
"""

第六步:进度报告
[3/71] AutoGLM 发布之后…
  解析完成: 95 段文字, 12 张图片
  ✓ 创建笔记: 501435173515
  ✓ 文字写入完成
  ✓ 图片插入: 12/12
  用时: 8.3s

进度:████████░░░░  42% (30/71)  预计剩余 ~12 分钟

转换脚本参考(需要查看实现时)

text
优先阅读这些文件来理解不同来源的导入逻辑:
- read_file("scripts/import_to_wps.py")  # 主导入流程、去重、图片插入
- read_file("scripts/scan_docs.py")      # 目录扫描、来源判断、筛选
- read_file("references/conversion-guide.md")  # 格式映射细节

笔记工具逐步写入(需要手动接管时)

python
# 1. 创建笔记
create_note(title="文档标题")

# 2. 获取初始 block ID
get_note_outline(note_id=note_id)

# 3. 写入内容(分批,每批 4 个 block)
batch_edit(note_id=note_id, operations=[
    {"op": "replace", "block_id": first_block_id, "content": "<h1>标题</h1>"},
    {"op": "insert", "anchor_id": first_block_id, "position": "after",
     "content": "<p>2025-04-21 18:30 | <tag>#推文</tag></p>"},
])

# 4. 插入图片(新版 WPS 支持后台插图)
insert_image(note_id=note_id, anchor_id=placeholder_block_id,
             position="before", src="data:image/jpeg;base64,...")

各来源特殊处理

微信公众号 HTML(原文.html)

微信公众号文章使用内联 CSS 样式表达富文本,需要从 HTML 解析格式:

目录结构:

文章目录/
  原文.html        ← 完整 HTML,含内联样式
  meta.json        ← {"title": "...", "publish_time": "2025-04-21 18:30", "url": "..."}
  图片/            ← 本地图片文件
    image_001.jpg
    image_002.jpg

HTML 解析要点:

  • 正文在 #js_content 容器内
  • 图片在 <img data-src="https://..."> 属性中(不是 src)
  • 标题通过 font-size >= 18px 的 span 推断
  • 颜色通过 span 的 color: rgb(...) 样式提取,需映射到 WPS 预设色
  • 粗体通过 font-weight: 700 或 bold 识别

内联样式 → WPS XML 映射:

python
# font-size >= 18px → <h2>
# font-weight: 700|bold → <strong>
# font-style: italic → <em>
# color: rgb(R,G,B) → <span fontColor="#WPS预设色">(需颜色映射)

详细颜色映射规则见 references/conversion-guide.md 第 10 节。


Show full SKILL.md (174 more words)Show less
Obsidian Vault
  1. Wiki 链接:[[文件名]] → 纯文本;[[文件名|显示名]] → 显示名
  2. 标签:#标签名 → <tag>#标签名</tag>
  3. Callouts:> [!NOTE] → WPS <highlightBlock>
  4. Frontmatter:YAML 头部提取为 meta 信息
  5. 图片路径:先找同目录 → 再找 attachments/ → 再找 Vault 根目录

思源笔记(SiYuan)
  • .sy 文件是 JSON 格式的 Block Tree
  • 图片路径格式为 assets/xxx.png,实际文件在 <工作空间>/data/assets/
  • 详细节点类型映射见 references/conversion-guide.md 第 6 节

故障排查

get_note_outline 只返回 100 个 block,后面内容丢失

使用翻页方案,见"WPS API 关键限制"第 1 条。

图片插入后点击显示加载失败

旧版 WPS(< 0.1.4)insert_image 的 fileID 绑定到前台笔记,升级到最新版本可解决。

insert_image 报 IMAGE_FETCH_FAILED
  • 检查 base64 data URI 格式:必须是 data:image/jpeg;base64,<数据>(不能是 data:application/octet-stream)
  • 大图片优先使用稳定 URL;如果只能传 base64,确保传入的是完整 data URI
anchor 失效导致内容截断

实现 4 次重试逻辑,每次失败后重新获取 last_block_id,见"WPS API 关键限制"第 4 条。

list_notes / search_notes 返回结果过多

缩小筛选范围或主动分页,见"WPS API 关键限制"第 3 条。

pandoc 未安装
bash
brew install pandoc  # macOS
pdfplumber 未安装
bash
pip3 install pdfplumber
PDF 扫描版无文字(需要 OCR)
bash
pip3 install pytesseract pdf2image
brew install tesseract

依赖项清单

工具用途安装
笔记 Agent 工具(create_note / get_note_outline / batch_edit / insert_image / read_note / read_blocks / search_notes / sync_note)WPS 笔记写入、去重与回读(必须)宿主提供
beautifulsoup4HTML 解析(公众号/网页)pip3 install beautifulsoup4
lxmlHTML 解析加速pip3 install lxml
pandocDOCX → Markdownbrew install pandoc
pdfplumberPDF 文本/表格提取pip3 install pdfplumber
pypdfPDF 图片提取备选pip3 install pypdf
markitdownPPTX → Markdownpip3 install "markitdown[pptx]"
pandasXLSX 读取pip3 install pandas openpyxl
pillow图片处理/base64转换pip3 install pillow
python-frontmatterYAML Frontmatter 解析pip3 install python-frontmatter

详细转换逻辑见 references/conversion-guide.md。

© wpsnote, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (scripts) in skills/doc-importer of wpsnote/wpsnote-skills.

  • SKILL.md
  • scripts/convert.py
  • scripts/import_to_wps.py
  • scripts/scan_docs.py

Open the folder on GitHubat commit 2d9885d

Compare with similar skills

Doc Importer next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Doc Importer compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Doc Importer this skillwpsnote/wpsnote-skills179—~2.7kAutomated safety check: PassMIT
MarkitdownImCa0/just-laws78114 repos~3.2kAutomated safety check: NotesMIT
Markitdownjimmc414/Kosmos5942 repos~1.7kAutomated safety check: PassNone
Office DocumentsZS520L/HanakoPro102—~1.6kAutomated safety check: PassApache-2.0
File Intelearlyaidopters/second-brain193—~481Automated safety check: PassNone
PaperJSX Document Generatorcomposio-community/awesome-codex-skills17k—~842Automated safety check: PassApache-2.0

Similar skills

  • Markitdown

    ImCa0/just-laws

    Convert files and office documents to Markdown. An agent skill from ImCa0/just-laws.

    781 GitHub starsUsed in 14 repos~3.2k tokens
    Documents & OfficeAuto-check: notes
  • Markitdown

    jimmc414/Kosmos

    Convert various file formats (PDF, Office documents, images, audio, web content, structured data) to Markdown optimized for LLM processing.

    594 GitHub starsUsed in 2 repos~1.7k tokens
    Documents & OfficeAuto-check passed
  • Office Documents

    ZS520L/HanakoPro

    A skill your agent uses when the user asks to open, read, inspect, understand, summarize, analyze, extract tables/text from, modify, update, repair, split, merge, rotate, or convert information from…

    102 GitHub stars~1.6k tokensUpdated 4 mo ago
    Documents & OfficeAuto-check passed
  • File Intel

    earlyaidopters/second-brain

    Run the Gemini file processor on any folder — extracts content from PDF, PPTX, XLSX, DOCX, CSV, JSON, and any text format, then generates Obsidian-ready summaries.

    193 GitHub stars~481 tokensUpdated 6 mo ago
    Documents & OfficeAuto-check passed
  • PaperJSX Document Generator

    composio-community/awesome-codex-skills

    Generates PPTX, DOCX, XLSX and PDF files from a JSON layout spec through PaperJSX's packages, creating new documents rather than editing them.

    17k GitHub stars~842 tokensUpdated 2 mo ago
    Documents & OfficeAuto-check passed
  • Skill Doc Delivery

    nyldn/claude-octopus

    Convert markdown to DOCX, PPTX, XLSX, PDF office documents — use when you need exportable deliverables

    4.2k GitHub starsUsed in 1 repo~2.4k tokens
    Documents & OfficeAuto-check passed

More from wpsnote/wpsnote-skills

All 37 skills in this repo
  • Skill Creator

    wpsnote/wpsnote-skills

    Create new skills or improve existing skills by clarifying intent, drafting SKILL.md instructions, organizing supporting resources, and doing lightweight manual review.

    179 GitHub stars~1.3k tokensUpdated 4 mo ago
    Auto-check passed
  • Content Digest

    wpsnote/wpsnote-skills

    将任意内容提炼为结构化知识笔记,自动保存到 WPS 笔记。只要用户给出任何内容(链接、图片、本地文件、粘贴文字)并有保存笔记的意图,就应使用此 skill。常见触发词:「总结」「提炼」「做笔记」「读书笔记」「学习笔记」「整理成笔记」「帮我看看」「帮我解读」「记一下」「存下来」「整理一下」「帮我归纳」。也适用于用户直接给出…

    179 GitHub stars~1.7k tokensUpdated 4 mo ago
    Auto-check passed
  • Content Creator

    wpsnote/wpsnote-skills

    【内容创作起点】用户想要写文章/公众号/文案时的首选skill. An agent skill from wpsnote/wpsnote-skills.

    179 GitHub stars~1.2k tokensUpdated 4 mo ago
    Auto-check passed
  • Image Gen

    wpsnote/wpsnote-skills

    AI 图像生成助手,支持文生图和图生图,对接 OpenRouter / 阿里云百炼 / 火山方舟 / Google Gemini。

    179 GitHub stars~1.3k tokensUpdated 4 mo ago
    Auto-check passed
  • Literature Reader

    wpsnote/wpsnote-skills

    阅读、分析并总结学术文献(PDF论文),生成结构化的文献概要笔记。核心能力:论文元信息提取、研究问题识别、方法论梳理、实验结果分析、个人评价生成,以及多篇文献横向对比。支持中英文论文、单篇精读与批量文献综述。当用户提供 PDF 论文文件、要求阅读文献、总结论文、文献综述、论文概要、论文精读、paper summary、paper review、读论文、读…

    179 GitHub stars~793 tokensUpdated 4 mo ago
    Auto-check passed
  • Paper Researcher

    wpsnote/wpsnote-skills

    学术论文全流程助手:搜索论文、下载 PDF、存入 WPS 笔记、精读分析。当用户说"搜论文"、"找论文"、"下载论文"、"读论文"、"帮我找 paper"、"搜一下 XXX 相关的论文"、"把这篇论文存到笔记"、"分析这篇论文"、"帮我做文献调研"时触发。支持 arXiv 和 OpenAlex 两个数据源,自动完成搜索→下载→转文本→写入 WPS 笔记→分析的完整闭环。

    179 GitHub stars~970 tokensUpdated 4 mo ago
    Auto-check passed

Questions about Doc Importer

What does Doc Importer do?

将本地文档批量导入到 WPS 笔记。支持扫描 Obsidian Vault、思源笔记、微信公众号存档、 下载目录或任意用户指定目录,自动识别 HTML、Markdown、PDF、DOCX、PPTX、XLSX 等格式, 转换为 WPS 笔记可读内容并保留图片和富文本格式。. Doc Importer is an agent skill from wpsnote/wpsnote-skills.

When should I use Doc Importer?

Doc Importer fits situations like: tasks that involve PowerPoint presentations; tasks that involve Word documents; tasks that involve Excel spreadsheets.

How do I install Doc Importer in Claude Code?

Run `npx skills add wpsnote/wpsnote-skills --skill doc-importer -a claude-code`. Or copy the skill folder (skills/doc-importer in wpsnote/wpsnote-skills) into .claude/skills/doc-importer in your project. Claude Code loads it when a task matches its description.

How do I install Doc Importer in Codex?

Run `npx skills add wpsnote/wpsnote-skills --skill doc-importer -a codex`. Or copy the skill folder (skills/doc-importer in wpsnote/wpsnote-skills) into .agents/skills/doc-importer in your project. Codex loads it when a task matches its description.

Can I use Doc Importer in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add wpsnote/wpsnote-skills --skill doc-importer -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/doc-importer, .gemini/skills/doc-importer, .github/skills/doc-importer and .opencode/skills/doc-importer in your project.

What does Doc Importer need to run?

Going by SKILL.md and its folder, Doc Importer needs Python for the scripts in its folder and the command-line tools its instructions call (pip3 and brew). Our summary lists: Python 3.

Does Doc Importer access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Doc Importer safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Doc Importer use?

Doc Importer is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Doc Importer use?

About 2.7k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Doc Importer?

Skills that share tags, products or a category with Doc Importer: Markitdown (ImCa0/just-laws, 781 stars), Markitdown (jimmc414/Kosmos, 594 stars), Office Documents (ZS520L/HanakoPro, 102 stars) and File Intel (earlyaidopters/second-brain, 193 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Doc Importer?

wpsnote (a GitHub organization) maintains it in wpsnote/wpsnote-skills, which has 179 GitHub stars. The repository holds 37 skills in this directory. The repository was last updated on May 25, 2026.

Source: wpsnote/wpsnote-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.