Markitdown
ImCa0/just-laws
Convert files and office documents to Markdown. An agent skill from ImCa0/just-laws.
将本地文档批量导入到 WPS 笔记。支持扫描 Obsidian Vault、思源笔记、微信公众号存档、 下载目录或任意用户指定目录,自动识别 HTML、Markdown、PDF、DOCX、PPTX、XLSX 等格式, 转换为 WPS 笔记可读内容并保留图片和富文本格式。
$ npx skills add wpsnote/wpsnote-skills --skill doc-importer -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install wpsnote/wpsnote-skills doc-importer --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/wpsnote/wpsnote-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/doc-importer .claude/skills/doc-importer && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "doc-importer" agent skill from https://github.com/wpsnote/wpsnote-skills/tree/main/skills/doc-importer into .claude/skills/doc-importer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "doc-importer", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/wpsnote/wpsnote-skills/tree/main/skills/doc-importerType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add wpsnote/wpsnote-skills --skill doc-importer -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install wpsnote/wpsnote-skills doc-importer --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/wpsnote/wpsnote-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/doc-importer .agents/skills/doc-importer && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "doc-importer" agent skill from https://github.com/wpsnote/wpsnote-skills/tree/main/skills/doc-importer into .agents/skills/doc-importer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "doc-importer", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add wpsnote/wpsnote-skills --skill doc-importer -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install wpsnote/wpsnote-skills doc-importer --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/wpsnote/wpsnote-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/doc-importer .cursor/skills/doc-importer && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "doc-importer" agent skill from https://github.com/wpsnote/wpsnote-skills/tree/main/skills/doc-importer into .cursor/skills/doc-importer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "doc-importer", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/wpsnote/wpsnote-skills.git --path skills/doc-importer--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add wpsnote/wpsnote-skills --skill doc-importer -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install wpsnote/wpsnote-skills doc-importer --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/wpsnote/wpsnote-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/doc-importer .gemini/skills/doc-importer && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "doc-importer" agent skill from https://github.com/wpsnote/wpsnote-skills/tree/main/skills/doc-importer into .gemini/skills/doc-importer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "doc-importer", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install wpsnote/wpsnote-skills doc-importerInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add wpsnote/wpsnote-skills --skill doc-importer -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/wpsnote/wpsnote-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/doc-importer .github/skills/doc-importer && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "doc-importer" agent skill from https://github.com/wpsnote/wpsnote-skills/tree/main/skills/doc-importer into .github/skills/doc-importer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "doc-importer", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add wpsnote/wpsnote-skills --skill doc-importer -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install wpsnote/wpsnote-skills doc-importer --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/wpsnote/wpsnote-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/doc-importer .opencode/skills/doc-importer && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "doc-importer" agent skill from https://github.com/wpsnote/wpsnote-skills/tree/main/skills/doc-importer into .opencode/skills/doc-importer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "doc-importer", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
doc-importer将本地文档批量导入到 WPS 笔记。支持扫描 Obsidian Vault、思源笔记、微信公众号存档、 下载目录或任意用户指定目录,自动识别 HTML、Markdown、PDF、DOCX、PPTX、XLSX 等格式, 转换为 WPS 笔记可读内容并保留图片和富文本格式。
Doc Importer is an agent skill from wpsnote/wpsnote-skills. 将本地文档批量导入到 WPS 笔记。支持扫描 Obsidian Vault、思源笔记、微信公众号存档、 下载目录或任意用户指定目录,自动识别 HTML、Markdown、PDF、DOCX、PPTX、XLSX 等格式, 转换为 WPS 笔记可读内容并保留图片和富文本格式。 当用户说「导入文档到 WPS 笔记」「把我的 Obsidian 笔记导入」「导入思源笔记」 「把下载的 PDF/Word/PPT 导入笔记」「把公众号文章导入笔记」「同步本地文档到 WPS 笔记」时触发。 不适用于直接编辑 WPS 笔记内容、文档格式转换(不导入到笔记)。
Its SKILL.md is about 2.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including scripts (for example `scripts/convert.py`, `scripts/import_to_wps.py` and `scripts/scan_docs.py`).
It sits in Documents & Office, covering PowerPoint presentations, Word documents and Excel spreadsheets. It works with Obsidian, Microsoft Excel, Microsoft PowerPoint and Microsoft Word. The licence is MIT.
5 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 2d9885d. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 3 files in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
pip3brewFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use pip3, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Doc Importer loads about 2.7k tokens when it runs. Until then it costs about 71 tokens; SKILL.md has 454 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from wpsnote/wpsnote-skills at commit 2d9885d, republished under its MIT licence (© wpsnote). 454 words, ~2,683 tokens.
.claude/skills/doc-importer/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.将本地文档(Obsidian、思源笔记、微信公众号 HTML、下载目录或任意目录)批量导入到 WPS 笔记。
当前技能的笔记写入统一走笔记 Agent 工具;如需理解批量扫描和转换逻辑,可先读取本 Skill 自带脚本:
参考实现:
- read_file("scripts/scan_docs.py") # 目录扫描与来源识别
- read_file("scripts/import_to_wps.py") # 转换与导入主流程
实际写入:
- create_note
- get_note_outline
- batch_edit
- insert_image
- read_note / read_blocks
- search_notes
- sync_note| 来源 | 说明 | 典型目录结构 |
|---|---|---|
| Obsidian Vault | 扫描 .md 文件,保留 wiki 链接、Callout、标签 | Vault/笔记名.md + attachments/ |
| 思源笔记 | 扫描 .sy 文件(JSON格式),提取 Block Tree | SiYuan/data/笔记本/文档.sy |
| 微信公众号 | 解析 原文.html(含富文本格式、内联样式) | 文章名/原文.html + 图片/ |
| 任意目录 | 用户指定路径,递归扫描子目录 | 任意 |
| 格式 | 转换方式 | 图片处理 | 富文本保留 |
|---|---|---|---|
.html | BeautifulSoup 解析内联样式 | 本地图片 base64 | ✅ 颜色/粗体/标题 |
.md / .markdown | 直接解析,转 WPS XML | 本地图片 base64 | 基本格式 |
.pdf | pdfplumber 提取文本 + pdfimages 提取图片 | 提取嵌入图片 | 标题推断 |
.docx | pandoc 转 markdown,提取 word/media/ | 解包提取 | 基本格式 |
.pptx | markitdown 提取文本 + 解包媒体 | 解包提取 | 幻灯片结构 |
.xlsx | pandas 读取,转 WPS table | 不含图片 | 表格结构 |
.txt | 直接读取 | 不含图片 | 无 |
.sy | JSON 解析思源 Block Tree | 提取 assets 图片 | 全部 |
在实际导入前,必须了解以下 WPS 笔记 API 的重要限制,否则会导致内容丢失或图片插入失败。
get_note_outline 默认只返回 100 个 blockget_note_outline 默认按分页视图返回 block,blocks 数组可能只覆盖前一部分内容,但 block_count 会返回真实总数。
影响:大文章(> 100 个 block)如果只靠首屏大纲查占位符,只能找到前一部分 block,后半段图片会丢失。
解决方案:先用 get_note_outline 拿首批 block,再用 read_blocks 续读:
def get_all_blocks(note_id):
"""翻页获取笔记全部 blocks"""
# 第一页用 get_note_outline(有 preview 字段)
r = get_note_outline(note_id=note_id)
data = r.get('data', {})
total = data.get('block_count', 0)
blocks = list(data.get('blocks', []))
last_id = blocks[-1]['id'] if blocks else None
# 超过首批范围时用 read_blocks 续读
while len(blocks) < total and last_id:
r2 = read_blocks(note_id=note_id, block_id=last_id, after=100)
new_blocks = (r2.get('data') or {}).get('blocks', [])
if not new_blocks:
break
blocks.extend(new_blocks)
last_id = new_blocks[-1]['id']
return blocks
get_note_outline返回的 block 常带preview字段;read_blocks更适合拿完整 XML 内容。
insert_image 在旧版 WPS 中只能插到前台笔记旧版本(< 0.1.4):insert_image 的图片 fileID 绑定到当前 WPS 客户端 UI 中打开的笔记,即使传了 note_id 参数,图片也会错误地关联到前台笔记。
新版本(>= 0.1.4):已修复,insert_image 可以直接后台插图到任意笔记,无需切换前台。
处理策略:
list_notes 返回大数据时要主动收敛范围list_notes 一次返回的数据过大时,不适合拿来做全文去重或全量扫描判断。
解决方案:
search_notes 按标题关键词搜索list_notes 当成唯一真相来源连续调用 batch_edit 插入内容时,WPS 内部可能重新索引,导致前一次操作返回的 anchor_id 失效。如果不处理,后续内容会静默跳过,造成文章后半段内容丢失。
必须实现重试逻辑:
def do_insert(note_id, content, get_anchor_fn, max_retries=4):
"""带重试的内容插入,自动刷新 anchor"""
anchor = get_anchor_fn()
for attempt in range(max_retries):
if attempt > 0:
time.sleep(1.5 * attempt)
anchor = get_anchor_fn() # 重新获取最新 anchor
res = batch_edit(note_id, [{
'op': 'insert', 'anchor_id': anchor,
'position': 'after', 'content': content
}])
if res.get('ok') is not False:
anchor = get_anchor_fn()
return True
return False # 4次都失败才放弃获取最新 anchor(最后一个 block):
def get_last_block_id(note_id):
r = get_note_outline(note_id=note_id)
blocks = (r.get('data') or {}).get('blocks', [])
return blocks[-1]['id'] if blocks else None大量内容写入时,BATCH_SIZE(每次 batch_edit 的 block 数)建议设为 4,过大会增加 anchor 失效概率。
BATCH_SIZE = 4 # 不要超过 8,否则容易出现 anchor 失效优先根据用户已给出的目录 / 文件线索判断来源类型,不足时再 ask_user 补最少必要信息:
优先识别这些常见来源线索:
- Obsidian Vault:`.md` 文件 + `attachments/` 或同类附件目录
- 思源笔记:`.sy` 文件 + `data/assets/`
- 微信公众号存档:`原文.html` + 图片目录 + `meta.json`
- 下载目录混合导入:PDF / DOCX / PPTX / XLSX / Markdown / HTML 混合存在如需对照扫描规则,用 `read_file("scripts/scan_docs.py")` 查看来源识别和输出字段约定;
整理后的清单至少要包含路径、标题、时间、格式、图片数和预计 block 数。输出 JSON 格式:
{
"source_type": "wechat_mp",
"root_path": "/Users/xxx/Documents/articles",
"files": [
{
"path": "/Users/xxx/Documents/articles/文章名/原文.html",
"rel_path": "文章名/原文.html",
"title": "文章标题",
"publish_time": "2025-04-21 18:30",
"size_bytes": 204800,
"modified": "2025-04-22T10:00:00",
"format": "html",
"estimated_images": 12,
"estimated_blocks": 180
}
],
"total": 71,
"formats": {"html": 71}
}
estimated_blocks帮助预判是否会超过 100 block,提前提示用户。
扫描到 71 个文件:
- HTML: 71 个
文件列表:
1. AutoGLM 发布之后,如今国产大模型终于长出了手。 (2025-03-31, 12张图)
2. 你可能看不懂扣子空间为什么重要… (2025-04-21, 17张图)
...(超过20个时截断,告知总数)
请问你想如何导入?
[A] 全部导入(71个文件)
[B] 手动选择(输入文件编号,如:1,3,5-10)
[S] 跳过已有标题的笔记(根据笔记标题去重)导入前通过标题检查 WPS 笔记中是否已存在同名笔记。
注意:不要把 list_notes 当成去重主入口,改用 search_notes 按标题关键词搜索:
def check_exists(title):
r = search_notes(keyword=title[:20], limit=5)
notes = (r.get('data') or {}).get('notes', [])
return next((n for n in notes if n['title'] == title), None)发现重复时询问用户:
发现以下笔记在 WPS 中已存在:
- 《AutoGLM 发布之后…》(最后更新:2025-05-01)
如何处理?
[O] 覆盖 [S] 跳过 [A] 追加 [RA] 对所有冲突应用相同策略推荐流程(以 HTML 富文本为例):
# 1. 解析 HTML,提取内容段落和图片
segments = html_to_segments(html_path, img_dir)
# segments 格式:[('xml', '<p>文字</p>'), ('img', Path('图片/image_001.jpg')), ...]
# 2. 创建笔记
note_id = create_note(title)
# 3. 写入标题行 + meta 行(时间、标签)
write_header(note_id, title, publish_time, tag)
# 4. 批量写入正文(图片先插占位符)
write_content_with_placeholders(note_id, segments)
# 5. 翻页查找所有占位符 block_id(用 get_all_blocks)
ph_map = find_placeholders(note_id)
# 6. 逐个插入真实图片,替换占位符
for idx, img_path in img_list:
insert_image(note_id, ph_map[idx]['block_id'], img_path)
delete_placeholder(note_id, ph_map[idx]['block_id'])Meta 行格式(根据来源类型灵活调整):
# 微信公众号:简洁的时间 + 标签
f'<p>{publish_time} | <tag id="{tag_id}">#推文</tag></p>'
# 通用文档:完整 meta blockquote
"""
<blockquote>
<p>📄 <strong>来源</strong>:{rel_path}</p>
<p>🕒 <strong>修改时间</strong>:{modified_time}</p>
<p>🔄 <strong>导入时间</strong>:{import_time}</p>
</blockquote>
"""[3/71] AutoGLM 发布之后…
解析完成: 95 段文字, 12 张图片
✓ 创建笔记: 501435173515
✓ 文字写入完成
✓ 图片插入: 12/12
用时: 8.3s
进度:████████░░░░ 42% (30/71) 预计剩余 ~12 分钟优先阅读这些文件来理解不同来源的导入逻辑:
- read_file("scripts/import_to_wps.py") # 主导入流程、去重、图片插入
- read_file("scripts/scan_docs.py") # 目录扫描、来源判断、筛选
- read_file("references/conversion-guide.md") # 格式映射细节# 1. 创建笔记
create_note(title="文档标题")
# 2. 获取初始 block ID
get_note_outline(note_id=note_id)
# 3. 写入内容(分批,每批 4 个 block)
batch_edit(note_id=note_id, operations=[
{"op": "replace", "block_id": first_block_id, "content": "<h1>标题</h1>"},
{"op": "insert", "anchor_id": first_block_id, "position": "after",
"content": "<p>2025-04-21 18:30 | <tag>#推文</tag></p>"},
])
# 4. 插入图片(新版 WPS 支持后台插图)
insert_image(note_id=note_id, anchor_id=placeholder_block_id,
position="before", src="data:image/jpeg;base64,...")原文.html)微信公众号文章使用内联 CSS 样式表达富文本,需要从 HTML 解析格式:
目录结构:
文章目录/
原文.html ← 完整 HTML,含内联样式
meta.json ← {"title": "...", "publish_time": "2025-04-21 18:30", "url": "..."}
图片/ ← 本地图片文件
image_001.jpg
image_002.jpgHTML 解析要点:
#js_content 容器内<img data-src="https://..."> 属性中(不是 src)font-size >= 18px 的 span 推断color: rgb(...) 样式提取,需映射到 WPS 预设色font-weight: 700 或 bold 识别内联样式 → WPS XML 映射:
# font-size >= 18px → <h2>
# font-weight: 700|bold → <strong>
# font-style: italic → <em>
# color: rgb(R,G,B) → <span fontColor="#WPS预设色">(需颜色映射)详细颜色映射规则见 references/conversion-guide.md 第 10 节。
[[文件名]] → 纯文本;[[文件名|显示名]] → 显示名#标签名 → <tag>#标签名</tag>> [!NOTE] → WPS <highlightBlock>attachments/ → 再找 Vault 根目录.sy 文件是 JSON 格式的 Block Treeassets/xxx.png,实际文件在 <工作空间>/data/assets/references/conversion-guide.md 第 6 节get_note_outline 只返回 100 个 block,后面内容丢失使用翻页方案,见"WPS API 关键限制"第 1 条。
旧版 WPS(< 0.1.4)insert_image 的 fileID 绑定到前台笔记,升级到最新版本可解决。
insert_image 报 IMAGE_FETCH_FAILEDdata:image/jpeg;base64,<数据>(不能是 data:application/octet-stream)实现 4 次重试逻辑,每次失败后重新获取 last_block_id,见"WPS API 关键限制"第 4 条。
list_notes / search_notes 返回结果过多缩小筛选范围或主动分页,见"WPS API 关键限制"第 3 条。
brew install pandoc # macOSpip3 install pdfplumberpip3 install pytesseract pdf2image
brew install tesseract| 工具 | 用途 | 安装 |
|---|---|---|
笔记 Agent 工具(create_note / get_note_outline / batch_edit / insert_image / read_note / read_blocks / search_notes / sync_note) | WPS 笔记写入、去重与回读(必须) | 宿主提供 |
beautifulsoup4 | HTML 解析(公众号/网页) | pip3 install beautifulsoup4 |
lxml | HTML 解析加速 | pip3 install lxml |
pandoc | DOCX → Markdown | brew install pandoc |
pdfplumber | PDF 文本/表格提取 | pip3 install pdfplumber |
pypdf | PDF 图片提取备选 | pip3 install pypdf |
markitdown | PPTX → Markdown | pip3 install "markitdown[pptx]" |
pandas | XLSX 读取 | pip3 install pandas openpyxl |
pillow | 图片处理/base64转换 | pip3 install pillow |
python-frontmatter | YAML Frontmatter 解析 | pip3 install python-frontmatter |
详细转换逻辑见 references/conversion-guide.md。
© wpsnote, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 3 other files (scripts) in skills/doc-importer of wpsnote/wpsnote-skills.
Open the folder on GitHubat commit 2d9885d
Doc Importer next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Doc Importer this skillwpsnote/wpsnote-skills | 179 | — | ~2.7k | Automated safety check: Pass | MIT | |
| MarkitdownImCa0/just-laws | 781 | 14 repos | ~3.2k | Automated safety check: Notes | MIT | |
| Markitdownjimmc414/Kosmos | 594 | 2 repos | ~1.7k | Automated safety check: Pass | None | |
| Office DocumentsZS520L/HanakoPro | 102 | — | ~1.6k | Automated safety check: Pass | Apache-2.0 | |
| File Intelearlyaidopters/second-brain | 193 | — | ~481 | Automated safety check: Pass | None | |
| PaperJSX Document Generatorcomposio-community/awesome-codex-skills | 17k | — | ~842 | Automated safety check: Pass | Apache-2.0 |
ImCa0/just-laws
Convert files and office documents to Markdown. An agent skill from ImCa0/just-laws.
jimmc414/Kosmos
Convert various file formats (PDF, Office documents, images, audio, web content, structured data) to Markdown optimized for LLM processing.
ZS520L/HanakoPro
A skill your agent uses when the user asks to open, read, inspect, understand, summarize, analyze, extract tables/text from, modify, update, repair, split, merge, rotate, or convert information from…
earlyaidopters/second-brain
Run the Gemini file processor on any folder — extracts content from PDF, PPTX, XLSX, DOCX, CSV, JSON, and any text format, then generates Obsidian-ready summaries.
composio-community/awesome-codex-skills
Generates PPTX, DOCX, XLSX and PDF files from a JSON layout spec through PaperJSX's packages, creating new documents rather than editing them.
nyldn/claude-octopus
Convert markdown to DOCX, PPTX, XLSX, PDF office documents — use when you need exportable deliverables
wpsnote/wpsnote-skills
Create new skills or improve existing skills by clarifying intent, drafting SKILL.md instructions, organizing supporting resources, and doing lightweight manual review.
wpsnote/wpsnote-skills
将任意内容提炼为结构化知识笔记,自动保存到 WPS 笔记。只要用户给出任何内容(链接、图片、本地文件、粘贴文字)并有保存笔记的意图,就应使用此 skill。常见触发词:「总结」「提炼」「做笔记」「读书笔记」「学习笔记」「整理成笔记」「帮我看看」「帮我解读」「记一下」「存下来」「整理一下」「帮我归纳」。也适用于用户直接给出…
wpsnote/wpsnote-skills
【内容创作起点】用户想要写文章/公众号/文案时的首选skill. An agent skill from wpsnote/wpsnote-skills.
wpsnote/wpsnote-skills
AI 图像生成助手,支持文生图和图生图,对接 OpenRouter / 阿里云百炼 / 火山方舟 / Google Gemini。
wpsnote/wpsnote-skills
阅读、分析并总结学术文献(PDF论文),生成结构化的文献概要笔记。核心能力:论文元信息提取、研究问题识别、方法论梳理、实验结果分析、个人评价生成,以及多篇文献横向对比。支持中英文论文、单篇精读与批量文献综述。当用户提供 PDF 论文文件、要求阅读文献、总结论文、文献综述、论文概要、论文精读、paper summary、paper review、读论文、读…
wpsnote/wpsnote-skills
学术论文全流程助手:搜索论文、下载 PDF、存入 WPS 笔记、精读分析。当用户说"搜论文"、"找论文"、"下载论文"、"读论文"、"帮我找 paper"、"搜一下 XXX 相关的论文"、"把这篇论文存到笔记"、"分析这篇论文"、"帮我做文献调研"时触发。支持 arXiv 和 OpenAlex 两个数据源,自动完成搜索→下载→转文本→写入 WPS 笔记→分析的完整闭环。
Categories
将本地文档批量导入到 WPS 笔记。支持扫描 Obsidian Vault、思源笔记、微信公众号存档、 下载目录或任意用户指定目录,自动识别 HTML、Markdown、PDF、DOCX、PPTX、XLSX 等格式, 转换为 WPS 笔记可读内容并保留图片和富文本格式。. Doc Importer is an agent skill from wpsnote/wpsnote-skills.
Doc Importer fits situations like: tasks that involve PowerPoint presentations; tasks that involve Word documents; tasks that involve Excel spreadsheets.
Run `npx skills add wpsnote/wpsnote-skills --skill doc-importer -a claude-code`. Or copy the skill folder (skills/doc-importer in wpsnote/wpsnote-skills) into .claude/skills/doc-importer in your project. Claude Code loads it when a task matches its description.
Run `npx skills add wpsnote/wpsnote-skills --skill doc-importer -a codex`. Or copy the skill folder (skills/doc-importer in wpsnote/wpsnote-skills) into .agents/skills/doc-importer in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add wpsnote/wpsnote-skills --skill doc-importer -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/doc-importer, .gemini/skills/doc-importer, .github/skills/doc-importer and .opencode/skills/doc-importer in your project.
Going by SKILL.md and its folder, Doc Importer needs Python for the scripts in its folder and the command-line tools its instructions call (pip3 and brew). Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Doc Importer is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.7k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Doc Importer: Markitdown (ImCa0/just-laws, 781 stars), Markitdown (jimmc414/Kosmos, 594 stars), Office Documents (ZS520L/HanakoPro, 102 stars) and File Intel (earlyaidopters/second-brain, 193 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
wpsnote (a GitHub organization) maintains it in wpsnote/wpsnote-skills, which has 179 GitHub stars. The repository holds 37 skills in this directory. The repository was last updated on May 25, 2026.
Source: wpsnote/wpsnote-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.